John W. Sias

dblp:35/276 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
0since 2021 · last 2006
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 3 first-authorSoftware engineering, systems software and programming languages · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
8 papers
Processor architecture and microarchitecture · 80% Memory systems · 18% Performance modeling and evaluation · 2%
Software engineering, system software, and programming languages
7 papers
Compilers and program optimization · 100%

Topics — the 19 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture
instruction-level parallelism
0.132006
Beating In-Order Stalls with "Flea-Flicker" Two-Pass Pipelining · IEEE Trans. Computers 2006
Field-testing IMPACT EPIC research results in Itanium 2 · ISCA 2004
Accurate and efficient predicate analysis with binary decision diagrams · MICRO 2000
Processor architecture and microarchitecture › instruction-level parallelism
predicated execution
0.142001
The Program Decision Logic Approach to Predicated Execution · ISCA 1999
Integrated Predicated and Speculative Execution in the IMPACT EPIC Architecture · ISCA 1998
Program decision logic optimization using predication and control speculation · Proc. IEEE 2001
Processor architecture and microarchitecture
latency hiding
0.112006
Beating In-Order Stalls with "Flea-Flicker" Two-Pass Pipelining · IEEE Trans. Computers 2006
Memory systems › memory access latency
load latency
0.112006
Beating In-Order Stalls with "Flea-Flicker" Two-Pass Pipelining · IEEE Trans. Computers 2006
Processor architecture and microarchitecture
pipelining
0.112006
Beating In-Order Stalls with "Flea-Flicker" Two-Pass Pipelining · IEEE Trans. Computers 2006
Compilers and program optimization
predicated execution
0.122000
Accurate and efficient predicate analysis with binary decision diagrams · MICRO 2000
The Program Decision Logic Approach to Predicated Execution · ISCA 1999
Compilers and program optimization › instruction scheduling
instruction-level parallelism
0.032001
Program decision logic optimization using predication and control speculation · Proc. IEEE 2001
The Program Decision Logic Approach to Predicated Execution · ISCA 1999
Integrated Predicated and Speculative Execution in the IMPACT EPIC Architecture · ISCA 1998
Memory systems › cache
cache miss tolerance
0.012003
Beating in-order stalls with "flea-flicker" two-pass pipelining · MICRO 2003
Processor architecture and microarchitecture › microprocessor design › processor core design
in-order core
0.012003
Beating in-order stalls with "flea-flicker" two-pass pipelining · MICRO 2003
Processor architecture and microarchitecture › pipelining
pipeline stall
0.012003
Beating in-order stalls with "flea-flicker" two-pass pipelining · MICRO 2003
Compilers and program optimization › compiler optimization
control flow optimization
0.012001
Program decision logic optimization using predication and control speculation · Proc. IEEE 2001
Compilers and program optimization
predicated compilation
0.012001
Program decision logic optimization using predication and control speculation · Proc. IEEE 2001
Compilers and program optimization › compiler analysis
predicate analysis
0.012000
Accurate and efficient predicate analysis with binary decision diagrams · MICRO 2000
Processor architecture and microarchitecture › instruction set architecture
EPIC architecture
0.011998
Integrated Predicated and Speculative Execution in the IMPACT EPIC Architecture · ISCA 1998
Processor architecture and microarchitecture
speculative execution
0.011998
Integrated Predicated and Speculative Execution in the IMPACT EPIC Architecture · ISCA 1998
Compilers and program optimization
instruction scheduling
0.012006
Beating In-Order Stalls with "Flea-Flicker" Two-Pass Pipelining · IEEE Trans. Computers 2006
Performance modeling and evaluation
benchmarking
0.012004
Field-testing IMPACT EPIC research results in Itanium 2 · ISCA 2004
Processor architecture and microarchitecture
instruction set architecture
0.012001
Program decision logic optimization using predication and control speculation · Proc. IEEE 2001
Processor architecture and microarchitecture › instruction-level parallelism
VLIW
0.012001
Enhancing loop buffering of media and telecommunications applications using low-overhead predication · MICRO 2001

Methods — techniques the papers use, named apart from their topics

instruction marking · 0.1reduced ordered binary decision diagram · 0.1profiling · 0.1in situ evaluation · 0.1loop transformation · 0.1if-conversion · 0.1logic synthesis · 0.0
YearPublicationVenuePosition
2006 Beating In-Order Stalls with "Flea-Flicker" Two-Pass Pipelining
abstract
While compilers have generally proven adept at planning useful static instruction-level parallelism for in-order microarchitectures, the efficient accommodation of unanticipateable latencies, like those of load instructions, remains a vexing problem. Traditional out-of-order execution hides some of these latencies, but repeats scheduling work already done by the compiler and adds additional pipeline overhead. Other techniques, such as prefetching and multithreading, can hide some anticipateable, long-latency misses, but not the shorter, more diffuse stalls due to difficult-to-anticipate, first or second-level misses. Our work proposes a microarchitectural technique, two-pass pipelining, whereby the program executes on two in-order back-end pipelines coupled by a queue. The "advance" pipeline often defers instructions dispatching with unready operands rather than stalling. The "backup" pipeline allows concurrent resolution of instructions deferred by the first pipeline allowing overlapping of useful "advanced" execution with miss resolution. An accompanying compiler technique and instruction marking further enhance the handling of miss latencies. Applying our technique to an Itanium 2-like design achieves a speedup of 1.38x in mcf, the most memory-intensive SPECint2000 benchmark, and an average of 1.12 x across other selected benchmarks, yielding between 32 percent and 67 percent of an idealized out-of-order design's speedup at a much lower design cost and complexity.
Ronald D. Barnes, John W. Sias, Erik M. Nystrom, Sanjay J. Patel, Nacho Navarro, Wen-Mei W. Hwu
IEEE Trans. Computers2
2004 Field-testing IMPACT EPIC research results in Itanium 2
abstract
Explicitly-Parallel Instruction Computing (EPIC) provides architectural features, including predication and explicit control speculation, intended to enhance the compiler's ability to expose instruction-level parallelism (ILP) in control-intensive programs. Aggressive structural transformations using these features, though described in the literature, have not yet been fully characterized in complete systems. Using the Intel Itanium 2 microprocessor, the SPECint2000 benchmarks and the IMPACT Compiler for IA-64, a research compiler competitive with the best commercial compilers on the platform, we provide an in situ evaluation of code generated using aggressive, EPIC-enabled techniques in a reality-constrained microarchitecture. Our work shows a 1.13 average speedup (up to 1.50) due to these compilation techniques, relative to traditionally-optimized code at the same inlining and pointer analysis levels, and a 1.55 speedup (up to 2.30) relative to GNU GCC, a solid traditional compiler. Detailed results show that the structural compilation approach provides benefits far beyond a decrease in branch misprediction penalties and that it both positively and negatively impacts instruction cache performance. We also demonstrate the increasing significance of runtime effects, such as data cache and TLB, in determining end performance and the interaction of these effects with control speculation.
John W. Sias, Sain-Zee Ueng, Geoff A. Kent, Ian M. Steiner, Erik M. Nystrom, Wen-Mei W. Hwu
ISCA1
2003 Beating in-order stalls with "flea-flicker" two-pass pipelining
abstract
Accommodating the uncertain latency of load instructions is one of the most vexing problems in in-order microarchitecture design and compiler development. Compilers can generate schedules with a high degree of instruction-level parallelism but cannot effectively accommodate unanticipated latencies; incorporating traditional out-of-order execution into the microarchitecture hides some of this latency but redundantly performs work done by the compiler and adds additional pipeline stages. Although effective techniques, such as prefetching and threading, have been proposed to deal with anticipable, long latency misses, the shorter, more diffuse stalls due to difficult-to-anticipate, first- or second-level misses are less easily hidden on in-order architectures. This paper addresses this problem by proposing a microarchitectural technique, referred to as two-pass pipelining, wherein the program executes on two in-order back-end pipelines coupled by a queue. The "advance" pipeline executes instructions greedily, without stalling on unanticipated latency dependences (executing independent instructions while otherwise blocking instructions are deferred). The "backup" pipeline allows concurrent resolution of instructions that were deferred in the other pipeline, resulting in the absorption of shorter misses and the overlap of longer ones. This paper argues that this design is both achievable and a good use of transistor resources and shows results indicating that it can deliver significant speedups for in-order processor designs.
Ronald D. Barnes, Erik M. Nystrom, John W. Sias, Sanjay J. Patel, Nacho Navarro, Wen-Mei W. Hwu
MICRO3
2001 Enhancing loop buffering of media and telecommunications applications using low-overhead predication
abstract
Media- and telecommunications-focused processors, increasingly designed as deeply pipelined, statically-scheduled VLIWs, rely on loop buffers for low-overhead execution of simple loops. Key loops containing control flow pose a substantial problem-full predication has a high encoding overhead, and partial predication techniques do not support if-conversion, the transformation of general acyclic control flow into predicated blocks. Using a set of significant media processing benchmarks, drawn from MediaBench and contemporary telecommunications standards, we explore a compromise approach. We demonstrate a compiler using if-conversion and specialized loop transformations to arrange for 70-99% of fetched operations to come from a simple, statically managed 256-instruction loop buffer, saving instruction fetch power and eliminating branch penalties. To complement this we introduce a "niche" form of predication specialized to permit general if-conversion with only a single bit in the encoding of each operation and to eliminate much of the hardware overhead of a predicate register-based approach.
John W. Sias, Hillery C. Hunter, Wen-Mei W. Hwu
MICRO1
2001 Program decision logic optimization using predication and control speculation
abstract
The mainstream arrival of predication, a means other than branching of selecting instructions for execution, has required compiler architects to reformulate fundamental analyses and transformations. Traditionally, the compiler has generated branches straightforwardly to implement control flow designed by the programmer and has then performed sophisticated "global" optimizations. to move and optimize code around them. In this model, the inherent tie between the control state of the program and the location of the single instruction pointer serialized run-time evaluation of control and limited the extent to which the compiler could optimize the control structure of the program (without extensive code replication). Predication provides a means of control independent of branches and instruction fetch location, freeing both compiler and architecture from these restrictions; effective compilation of predicated code, however requires sophisticated understanding of the program's control structure. This paper explores a representational technique which, through direct code analysis, maps the program's control component into a canonical database, a reduced ordered binary decision diagram (ROBDD), which fully enables the compiler to utilize and manipulate predication. This abstraction is then applied to optimize the program's control component, transforming it into a form more amenable to instruction level parallel (ILP) execution.
Wen-Mei W. Hwu, David I. August, John W. Sias
Proc. IEEE3
2000 Accurate and efficient predicate analysis with binary decision diagrams
abstract
Functionality and performance of EPIC architectural features depend on extensive compiler support. Predication, one of these features, promises to reduce control flow overhead and to enhance optimization, provided that compilers can utilize it effectively. Previous work has established the need for accurate, direct predicate analysis and has demonstrated a few useful techniques, but has not provided an efficient, general framework. The paper presents the Predicate Analysis System (PAS), which maps knowledge of predicate and condition relations in general control flow onto a convenient logical substrate, the reduced ordered binary decision diagram. PAS is the first such framework to demonstrate direct, accurate, and efficient analysis of arbitrary condition and predicate define networks in arbitrary control flow.
John W. Sias, Wen-Mei W. Hwu, David I. August
MICRO1
1999 The Program Decision Logic Approach to Predicated Execution
abstract
Modern compilers must expose sufficient amounts of Instruction-Level Parallelism (ILP) to achieve the promised performance increases of superscalar and VLIW processors. One of the major impediments to achieving this goal has been inefficient programmatic control flow. Historically, the compiler has translated the programmer's original control structure directly into assembly code with conditional branch instructions. Eliminating inefficiencies in handling branch instructions and exploiting ILP has been the subject of much research. However, traditional branch handling techniques cannot significantly alter the program's inherent control structure. The advent of predication as a program control representation has enabled compilers to manipulate control in a form more closely related to the underlying program logic. This work takes full advantage of the predication paradigm by abstracting the program control flow into a logical form referred to as a program decision logic network. This network is modeled as a Boolean equation and minimized using modified versions of logic synthesis techniques. After minimization, the more efficient version of the program's original control flow is re-expressed in predicated code. Furthermore, this paper proposes extensions to the HPL PlayDoh predication model in support of more effective predicate decision logic network minimization. Finally, this paper shows the ability of the mechanisms presented to overcome limits on ILP previously imposed by rigid program control structure.
David I. August, John W. Sias, Jean-Michel Puiatti, Scott A. Mahlke, Daniel A. Connors, Kevin M. Crozier, Wen-Mei W. Hwu
ISCA2
1998 Integrated Predicated and Speculative Execution in the IMPACT EPIC Architecture
abstract
Explicitly Parallel Instruction Computing (EPIC) architectures require the compiler to express program instruction level parallelism directly to the hardware. EPIC techniques which enable the compiler to represent control speculation, data dependence speculation, and predication have individually been shown to be very effective. However these techniques have not been studied in combination with each other. This paper presents the IMPACT EPIC Architecture to address the issues involved in designing processors based on these EPIC concepts. In particular we focus on new execution and recovery models in which microarchitectural support for predicated execution is also used to enable efficient recovery from exceptions caused by speculatively executed instructions. This paper demonstrates that a coherent framework to integrate the three techniques can be elegantly designed to achieve much better performance than each individual technique could alone provide.
David I. August, Daniel A. Connors, Scott A. Mahlke, John W. Sias, Kevin M. Crozier, Ben-Chung Cheng, Patrick R. Eaton, Qudus B. Olaniran, Wen-Mei W. Hwu
ISCA4