E. Christopher Lewis

dblp:58/5295 · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
0since 2021 · last 2008
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6Software engineering, systems software and programming languages · 6 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
6 papers
Parallel and multicore computing · 53% Processor architecture and microarchitecture · 25% Memory systems · 14%
Software engineering, system software, and programming languages
5 papers
Compilers and program optimization · 61% Program analysis · 27% Operating systems · 12%
Network and information security
2 papers
Systems and software security · 100%

Topics — the 19 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Systems and software security
operating system security
0.112008
Overshadow: a virtualization-based approach to retrofitting protection in commodity operating systems · ASPLOS 2008
Systems and software security › virtualization security
virtualization-based security
0.112008
Overshadow: a virtualization-based approach to retrofitting protection in commodity operating systems · ASPLOS 2008
Parallel and multicore computing › parallel programming models
multithreaded programming
0.112007
Making the fast case common and the uncommon case simple in unbounded transactional memory · ISCA 2007
Parallel and multicore computing
transactional memory
0.112007
Making the fast case common and the uncommon case simple in unbounded transactional memory · ISCA 2007
Program analysis › dynamic analysis
dynamic instrumentation
0.112005
Low-Overhead Interactive Debugging via Dynamic Instrumentation with DISE · HPCA 2005
Processor architecture and microarchitecture › instruction set architecture
instruction set customization
0.012003
DISE: A Programmable Macro Engine for Customizing Applications · ISCA 2003
Compilers and program optimization
parallelizing compiler
0.012000
ZPL: A Machine Independent Programming Language for Parallel Computers · IEEE Trans. Software Eng. 2000
Parallel and multicore computing
parallel programming models
0.012000
ZPL: A Machine Independent Programming Language for Parallel Computers · IEEE Trans. Software Eng. 2000
Operating systems
commodity operating systems
0.012008
Overshadow: a virtualization-based approach to retrofitting protection in commodity operating systems · ASPLOS 2008
Memory systems › memory management
virtual memory
0.012008
Overshadow: a virtualization-based approach to retrofitting protection in commodity operating systems · ASPLOS 2008
Memory systems
cache coherence
0.012007
Making the fast case common and the uncommon case simple in unbounded transactional memory · ISCA 2007
Compilers and program optimization › loop optimization
array contraction
0.011998
The Implementation and Evaluation of Fusion and Contraction in Array Languages · PLDI 1998
Compilers and program optimization › domain-specific compilation
array language compilation
0.011998
The Implementation and Evaluation of Fusion and Contraction in Array Languages · PLDI 1998
Compilers and program optimization › loop transformation
loop fusion
0.011998
The Implementation and Evaluation of Fusion and Contraction in Array Languages · PLDI 1998
Compilers and program optimization
loop transformation
0.011998
The Implementation and Evaluation of Fusion and Contraction in Array Languages · PLDI 1998
Reconfigurable computing and FPGAs
programmable devices
0.012005
Low-Overhead Interactive Debugging via Dynamic Instrumentation with DISE · HPCA 2005
Distributed systems
communication optimization
0.011998
The Implementation and Evaluation of Fusion and Contraction in Array Languages · PLDI 1998
Parallel and multicore computing
data parallelism
0.011998
The Implementation and Evaluation of Fusion and Contraction in Array Languages · PLDI 1998
High-performance computing
scientific computing
0.011998
The Implementation and Evaluation of Fusion and Contraction in Array Languages · PLDI 1998

Methods — techniques the papers use, named apart from their topics

virtual machine based isolation · 0.2multi-shadowing · 0.2simulation · 0.1dynamic instruction macro-expansion · 0.1cycle-level simulation · 0.1hardware design · 0.1performance model · 0.1abstract parallel machine · 0.1loop fusion · 0.0array contraction · 0.0scalarization · 0.0
YearPublicationVenuePosition
2008 Overshadow: a virtualization-based approach to retrofitting protection in commodity operating systems
abstract
Commodity operating systems entrusted with securing sensitive data are remarkably large and complex, and consequently, frequently prone to compromise. To address this limitation, we introduce a virtual-machine-based system called Overshadow that protects the privacy and integrity of application data, even in the event of a total OScompromise. Overshadow presents an application with a normal view of its resources, but the OS with an encrypted view. This allows the operating system to carry out the complex task of managing an application's resources, without allowing it to read or modify them. Thus, Overshadow offers a last line of defense for application data.Overshadow builds on multi-shadowing, a novel mechanism that presents different views of physical memory, depending on the context performing the access. This primitive offers an additional dimension of protection beyond the hierarchical protection domains implemented by traditional operating systems and processor architectures.We present the design and implementation of Overshadow and show how its new protection semantics can be integrated with existing systems. Our design has been fully implemented and used to protect a wide range of unmodified legacy applications running on an unmodified Linux operating system. We evaluate the performance of our implementation, demonstrating that this approach is practical.
Tal Garfinkel, E. Christopher Lewis, Pratap Subrahmanyam, Carl A. Waldspurger, Dan Boneh, Jeffrey S. Dwoskin, Dan R. K. Ports
ASPLOS3
2008 Bantam: a customizable, java-based, classroom compiler
abstract
This paper introduces the Bantam Java compiler project, a new language and compiler designed specifically for the classroom Bantam Java, the source programming language, is a small subset of the Java language, which is a commonly-used language in introductory programming courses. Because Bantam Java is similar to Java, it leverages the student's existing intuition and the student can automatically apply what they learn in the course directly to Java. The Bantam Java project is also customizable (it supports several tools and targets), which gives instructors flexibility in designing course assignments. Finally, the Bantam Java compiler project includes a free, comprehensive, student manual which can be used in conjunction with any compiler textbook.
Marc L. Corliss, E. Christopher Lewis
SIGCSE2
2007 Making the fast case common and the uncommon case simple in unbounded transactional memory
abstract
Hardware transactional memory has great potential to simplify the creation ofcorrect and efficient multithreaded programs, allowing programmers to exploitmore effectively the soon-to-be-ubiquitous multi-core designs. Several recentproposals have extended the original bounded transactional memory to unboundedtransactional memory, a crucial step toward transactions becoming ageneral-purpose primitive. Unfortunately, supporting the concurrent executionof an unbounded number of unbounded transactions is challenging, and as aresult, many proposed implementations are complex.
Colin Blundell, Joseph Devietti, E. Christopher Lewis, Milo M. K. Martin
ISCA3
2005 Low-Overhead Interactive Debugging via Dynamic Instrumentation with DISE
abstract
Breakpoints, watchpoints, and conditional variants of both are essential debugging primitives, but their natural implementations often degrade performance significantly. Slowdown arises because the debugger - the tool implementing the breakpoint/watchpoint interface - is implemented in a process separate from the debugged application. Since the debugger evaluates the watchpoint expressions and conditional predicates to determine whether to invoke the user, a debugging session typically requires many expensive application-debugger context switches, resulting in slowdowns of 40,000 times or more in current commercial and open-source debuggers! In this paper, we present an effective and efficient implementation of (conditional) breakpoints and watchpoints that uses DISE to dynamically embed debugger logic into the running application. DISE (dynamic instruction stream editing) is a previously proposed, programmable hardware facility for dynamically customizing applications by transforming the instruction stream as it is decoded. DISE embedding preserves the logical separation of application and debugger nstructions are added dynamically and transparently, existing application code and data are not statically modified - and has little startup cost. Cycle-level simulation on the SPEC 2000 integer benchmarks shows that the DISE approach eliminates all unnecessary context switching, typically limits debugging overhead to 25% or less for a wide range of watch-points, and outperforms alternative implementations.
Marc L. Corliss, E. Christopher Lewis, Amir Roth
HPCA2
2005 The implementation and evaluation of dynamic code decompression using DISE
abstract
Code compression coupled with dynamic decompression is an important technique for both embedded and general-purpose microprocessors. Postfetch decompression , in which decompression is performed after the compressed instructions have been fetched, allows the instruction cache to store compressed code but requires a highly efficient decompression implementation. We propose implementing postfetch decompression using a new hardware facility called dynamic instruction stream editing (DISE). DISE provides a programmable decoder---similar in structure to those in many IA-32 processors---that is used to add functionality to an application by injecting custom code snippets into its fetched instruction stream. We present a DISE-based implementation of postfetch decompression and show that it naturally supports customized program-specific decompression dictionaries, enables parameterized decompression allowing similar-but-not-identical instruction sequences to share dictionary entries, and uses no decompression-specific hardware. We present extensive experimental results showing the virtue of this approach and evaluating the factors that impact its efficacy. We also present implementation-neutral results that give insight into the characteristics of any postfetch decompression technique. Our experiments not only demonstrate significant reduction in code size (up to 35%) but also significant improvements in performance (up to 20%) and energy (up to 10%).
Marc L. Corliss, E. Christopher Lewis, Amir Roth
ACM Trans. Embed. Comput. Syst.2
2003 DISE: A Programmable Macro Engine for Customizing Applications
abstract
Dynamic instruction stream editing (DISE) is a cooperative software-hardware scheme for efficiently adding customization functionality $e.g, safety/security checking, profiling, dynamic code decompression, and dynamic optimization - to an application. In DISE, application customization functions (ACFs) are formulated as rules for macro-expanding certain instructions into parameterized instruction sequences. The processor executes the rules on the fetched instructions, feeding the execution engine an instruction stream that contains ACF code. Dynamic instruction macro-expansion is widely used in many of today's processors to convert a complex ISA to an easier-to-execute, finer-grained internal form. DISE coopts this technology and adds a programming interface to it. DISE unifies the implementation of a large class of ACFs that would otherwise require either special-purpose hardware widgets or static binary rewriting. We show DISE implementations of two ACFs - memory fault isolation and dynamic code decompression - and their composition. Simulation shows that DISE ACFs have better performance than their software counterparts, and more flexibility (which sometimes translates into performance) than hardware implementations.
Marc L. Corliss, E. Christopher Lewis, Amir Roth
ISCA2
2003 A DISE implementation of dynamic code decompression
abstract
Code compression coupled with dynamic decompression is an important technique for both embedded and general-purpose microprocessors. Post-fetch decompression, in which decompression is performed after the compressed instructions have been fetched, allows the instruction cache to store compressed code but requires a highly efficient decompression implementation. We propose implementing post-fetch decompression using dynamic instruction stream editing (DISE), a programmable decoder---similar in structure to those in many IA32 processors---that is used to add functionality to an application by injecting custom code snippets into its fetched instruction stream. A DISE implementation of post-fetch decompression naturally supports customized program-specific decompression dictionaries, enables parameterized decompression allowing similar instruction sequences to share dictionary entries, and uses no decompression-specific hardware. Cycle-level simulation of DISE decompression shows that it can reduce static program size by 35% and execution time by 20%. Parameterized decompression, a feature unique to DISE, accounts for 20% of the code size reduction by making more effective use of the dictionary and allowing PC-relative branches to be included in compressed sequences. DISE-based compression can reduce total energy consumption by 10% and the energy-delay product by as much as 20%.
Marc L. Corliss, E. Christopher Lewis, Amir Roth
LCTES2
2000 A study of common pitfalls in simple multi-threaded programs
abstract
It is generally acknowledged that developing correct multi-threaded codes is difficult, because threads may interact with each other in unpredictable ways. The goal of this work is to discover common multi-threaded programming pitfalls, the knowledge of which will be useful in instructing new programmers and in developing tools to aid in multi-threaded programming. To this end, we study multi-threaded applications written by students from introductory operating systems courses. Although the applications are simple, careful inspection and the use of an automatic race detection tool reveal a surprising quantity and variety of synchronization errors. We describe and discuss these errors, evaluate the role of automated tools, and propose new tools for use in the instruction of multi-threaded programming.
Sung-Eun Choi, E. Christopher Lewis
SIGCSE2
2000 ZPL: A Machine Independent Programming Language for Parallel Computers
abstract
The goal of producing architecture-independent parallel programs is complicated by the competing need for high performance. The ZPL programming language achieves both goals by building upon an abstract parallel machine and by providing programming constructs that allow the programmer to "see" this underlying machine. This paper describes ZPL and provides a comprehensive evaluation of the language with respect to its goals of performance, portability, and programming convenience. In particular, we describe ZPt's machine-independent performance model, describe the programming benefits of ZPL's region-based constructs, summarize the compilation benefits of the language's high-level semantics, and summarize empirical evidence that ZPL has achieved both high performance and portability on diverse machines such as the IBM SP-2, Cray T3E, and SGI Power Challenge.
Bradford L. Chamberlain, Sung-Eun Choi, E. Christopher Lewis, Calvin Lin, Lawrence Snyder 0001, Derrick Weathersby
IEEE Trans. Software Eng.3
1999 Problem space promotion and its evaluation as a technique for efficient parallel computation
abstract
In this paper we describe a parallelprogrammingparadigm called problem space promotion (PSP), a technique that increases parallelism by reducing communication and synchronization.We present four algorithms that exploit PSP and evaluate their communication characteristics relative non-PSP solutions.Our analysis is aided by the use ofparallel algorithm notation that is concise, yet accurately reflects parallelism and communication costs.Our analysis illustrates circumstances under which the use of PSP is benejcial and detrimental to performance, and experiments on the Cray T3E attest to the validity of the analysis.We find that PSP can signtficantly improve the per$ormance and scaling behavior of certain computations, even when compared to existing high quality parallel algorithms.
Bradford L. Chamberlain, E. Christopher Lewis, Lawrence Snyder 0001
International Conference on Supercomputing2
1998 ZPL's WYSIWYG Performance Model
abstract
ZPL is a parallel array language designed for high performance scientific and engineering computations. Unlike other parallel languages, ZPL is founded on a machine model (the CTA) that accurately abstracts contemporary MIMD parallel computers. This makes it possible to correlate ZPL programs with machine behavior. As a result, programmers can reason about how code will perform on a typical parallel machine and thereby make informed decisions between alternative programming solutions. The paper describes ZPL's performance model and its syntactic cues for conveying operation cost. The what you see is what you get (WYSIWYG) nature of ZPL operations is demonstrated on the IBM SP-2, Intel Paragon, SGI Power Challenge, and Cray T3E. Additionally, the model is used to evaluate two algorithms for matrix multiplication. Experiments show that the performance model correctly predicts the faster solution on all four platforms for a range of problem sizes.
Bradford L. Chamberlain, Sung-Eun Choi, E. Christopher Lewis, Calvin Lin, Lawrence Snyder 0001, Derrick Weathersby
HIPS3
1998 The Implementation and Evaluation of Fusion and Contraction in Array Languages
abstract
Array languages such as Fortran 90, HPF and ZPL have many benefits in simplifying array-based computations and expressing data parallelism. However, they can suffer large performance penalties because they introduce intermediate arrays---both at the source level and during the compilation process---which increase memory usage and pollute the cache. Most compilers address this problem by simply scalarizing the array language and relying on a scalar language compiler to perform loop fusion and array contraction. We instead show that there are advantages to performing a form of loop fusion and array contraction at the array level. This paper describes this approach and explains its advantages. Experimental results show that our scheme typically yields runtime improvements of greater than 20% and sometimes up to 400%. In addition, it yields superior memory use when compared against commercial compilers and exhibits comparable memory use when compared with scalar languages. We also explore the interaction between these transformations and communication optimizations.
E. Christopher Lewis, Calvin Lin, Lawrence Snyder 0001
PLDI1