VLDB 2026 Research / reviewers in the wild / expert
William Y. Chen
dblp:36/5838
· DBLP profile ↗
18ranked-venue papers
5as first author
0since 2021 · last 1995
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 5 first-authorSoftware engineering, systems software and programming languages · 5
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
13 papers |
Processor architecture and microarchitecture · 84% Memory systems · 15% Parallel and multicore computing · 1% | |
| Software engineering, system software, and programming languages
10 papers |
Compilers and program optimization · 100% |
Topics — the 26 heaviest of 29, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Processor architecture and microarchitecture
instruction-level parallelism |
0.1 | 6 | 1995 | Three Architecutral Models for Compiler-Controlled Speculative Execution · IEEE Trans. Computers 1995 Dynamic Memory Disambiguation Using the Memory Conflict Buffer · ASPLOS 1994 Sentinel Scheduling for VLIW and Superscalar Processors · ACM Trans. Comput. Syst. 1993 |
Compilers and program optimization
instruction scheduling |
0.0 | 4 | 1995 | Three Architecutral Models for Compiler-Controlled Speculative Execution · IEEE Trans. Computers 1995 The Importance of Prepass Code Scheduling for Superscalar and Superpipelined Processors · IEEE Trans. Computers 1995 IMPACT: An Architectural Framework for Multiple-Instruction-Issue Processors · ISCA 1991 |
Processor architecture and microarchitecture
speculative execution |
0.0 | 3 | 1995 | Three Architecutral Models for Compiler-Controlled Speculative Execution · IEEE Trans. Computers 1995 Sentinel Scheduling for VLIW and Superscalar Processors · ACM Trans. Comput. Syst. 1993 Sentinel Scheduling for VLIW and Superscalar Processors · ASPLOS 1992 |
Processor architecture and microarchitecture
superscalar processor |
0.0 | 4 | 1995 | The Importance of Prepass Code Scheduling for Superscalar and Superpipelined Processors · IEEE Trans. Computers 1995 Compiler Code Transformations for Superscalar-Based High Performance Systems · SC 1992 Data Access Microarchitectures for Superscalar Processors with Compiler-Assisted Data Prefetching · MICRO 1991 |
Processor architecture and microarchitecture › instruction-level parallelism
compiler-controlled speculative execution |
0.0 | 2 | 1993 | Sentinel Scheduling for VLIW and Superscalar Processors · ACM Trans. Comput. Syst. 1993 Sentinel Scheduling for VLIW and Superscalar Processors · ASPLOS 1992 |
Compilers and program optimization › instruction scheduling
global instruction scheduling |
0.0 | 1 | 1995 | The Importance of Prepass Code Scheduling for Superscalar and Superpipelined Processors · IEEE Trans. Computers 1995 |
Compilers and program optimization › instruction scheduling
superblock scheduling |
0.0 | 1 | 1995 | Three Architecutral Models for Compiler-Controlled Speculative Execution · IEEE Trans. Computers 1995 |
Processor architecture and microarchitecture
instruction scheduling |
0.0 | 1 | 1995 | The Importance of Prepass Code Scheduling for Superscalar and Superpipelined Processors · IEEE Trans. Computers 1995 |
Processor architecture and microarchitecture › instruction-level parallelism
superscalar and VLIW processors |
0.0 | 2 | 1992 | Sentinel Scheduling for VLIW and Superscalar Processors · ASPLOS 1992 Effective compiler support for predicated execution using the hyperblock · MICRO 1992 |
Compilers and program optimization › interprocedural optimization
inlining |
0.0 | 1 | 1993 | The Effect of Code Expanding Optimizations on Instruction Cache Design · IEEE Trans. Computers 1993 |
Processor architecture and microarchitecture
instruction set architecture |
0.0 | 1 | 1993 | Register Connection: A New Approach to Adding Registers into Instruction Set Architectures · ISCA 1993 |
Processor architecture and microarchitecture
register file |
0.0 | 1 | 1993 | Register Connection: A New Approach to Adding Registers into Instruction Set Architectures · ISCA 1993 |
Compilers and program optimization › program transformation
compiler transformations |
0.0 | 1 | 1992 | Compiler Code Transformations for Superscalar-Based High Performance Systems · SC 1992 |
Compilers and program optimization › instruction scheduling
instruction-level parallelism |
0.0 | 1 | 1992 | Compiler Code Transformations for Superscalar-Based High Performance Systems · SC 1992 |
Compilers and program optimization
predicated execution |
0.0 | 1 | 1992 | Effective compiler support for predicated execution using the hyperblock · MICRO 1992 |
Processor architecture and microarchitecture › instruction-level parallelism
VLIW |
0.0 | 1 | 1992 | Compiler Code Transformations for Superscalar-Based High Performance Systems · SC 1992 |
Memory systems
cache |
0.0 | 1 | 1991 | Data Access Microarchitectures for Superscalar Processors with Compiler-Assisted Data Prefetching · MICRO 1991 |
Processor architecture and microarchitecture › instruction-level parallelism › VLIW
VLIW processor |
0.0 | 1 | 1991 | IMPACT: An Architectural Framework for Multiple-Instruction-Issue Processors · ISCA 1991 |
Processor architecture and microarchitecture › superscalar processor
wide-issue processor |
0.0 | 1 | 1991 | IMPACT: An Architectural Framework for Multiple-Instruction-Issue Processors · ISCA 1991 |
Memory systems
cache design |
0.0 | 2 | 1993 | The Effect of Code Expanding Optimizations on Instruction Cache Design · IEEE Trans. Computers 1993 An efficient architecture for loop based data preloading · MICRO 1992 |
Compilers and program optimization
register allocation |
0.0 | 1 | 1993 | Register Connection: A New Approach to Adding Registers into Instruction Set Architectures · ISCA 1993 |
Processor architecture and microarchitecture
exception handling |
0.0 | 1 | 1993 | Sentinel Scheduling for VLIW and Superscalar Processors · ACM Trans. Comput. Syst. 1993 |
Memory systems › cache › CPU cache
instruction cache |
0.0 | 1 | 1993 | The Effect of Code Expanding Optimizations on Instruction Cache Design · IEEE Trans. Computers 1993 |
Compilers and program optimization › instruction scheduling
compile-time scheduling |
0.0 | 1 | 1992 | Sentinel Scheduling for VLIW and Superscalar Processors · ASPLOS 1992 |
Parallel and multicore computing › loop transformation
loop parallelization |
0.0 | 1 | 1992 | Compiler Code Transformations for Superscalar-Based High Performance Systems · SC 1992 |
Memory systems › cache management › cache interference
cache pollution |
0.0 | 1 | 1991 | Data Access Microarchitectures for Superscalar Processors with Compiler-Assisted Data Prefetching · MICRO 1991 |
Methods — techniques the papers use, named apart from their topics
profile-based compilation · 0.0compiler-controlled speculation · 0.0compile-time scheduling · 0.0hardware memory conflict buffer · 0.0compiler repair code · 0.0performance evaluation · 0.0execution-driven simulation · 0.0speculative execution · 0.0loop unrolling · 0.0DOALL/DOACROSS transformations · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 1995 | The Importance of Prepass Code Scheduling for Superscalar and Superpipelined ProcessorsabstractSuperscalar and superpipelined processors utilize parallelism to achieve peak performance that can be several times higher than that of conventional scalar processors. In order for this potential to be translated into the speedup of real program, the compiler must be able to schedule instructions so that the parallel hardware is effectively utilized. Previous work has shown that prepass code scheduling helps to produce a better schedule for scientific programs, but the importance of prescheduling has never been demonstrated for control-intensive non-numeric programs. These programs are significantly different from the scientific programs because they contain frequent branches. The compiler must do global scheduling in order to find enough independent instructions. In this paper, the code optimizer and scheduler of the IMPACT-I C compiler is described. Within this framework, we study the importance of prepass code scheduling for a set of production C programs. It is shown that, in contrast to the results previously obtained for scientific programs, prescheduling is not important for compiling control-intensive programs to the current generation of superscalar and superpipelined processors. However, if some of the current restrictions on upward code motion can be removed in future architectures, prescheduling would substantially improve the execution time of this class of programs on both superscalar and superpipelined processors.> Pohua P. Chang, Daniel M. Lavery, Scott A. Mahlke, William Y. Chen, Wen-Mei W. Hwu |
IEEE Trans. Computers | 4 |
| 1995 | Three Architecutral Models for Compiler-Controlled Speculative ExecutionabstractTo effectively exploit instruction level parallelism, the compiler must move instructions across branches. When an instruction is moved above a branch that it is control dependent on, it is considered to be speculatively executed since it is executed before it is known whether or not its result is needed. There are potential hazards when speculatively executing instructions. If these hazards can be eliminated, the compiler can more aggressively schedule the code. The hazards of speculative execution are outlined in this paper. Three architectural models: restricted, general, and boosting, which have increasing amounts of support for removing these hazards are discussed. The performance gained by each level of additional hardware support is analyzed using the IMPACT C compiler which performs superblock scheduling for superscalar and superpipelined processors.> Pohua P. Chang, Nancy J. Warter, Scott A. Mahlke, William Y. Chen, Wen-Mei W. Hwu |
IEEE Trans. Computers | 4 |
| 1994 | Dynamic Memory Disambiguation Using the Memory Conflict BufferabstractTo exploit instruction level parallelism, compilers for VLIW and superscalar processors often employ static code scheduling. However, the available code reordering may be severely restricted due to ambiguous dependences between memory instructions. This paper introduces a simple hardware mechanism, referred to as the memory conflict buffer, which facilitates static code scheduling in the presence of memory store/load dependences. Correct program execution is ensured by the memory conflict buffer and repair code provided by the compiler. With this addition, significant speedup over an aggressive code scheduling model can be achieved for both non-numerical and numerical programs. David M. Gallagher, William Y. Chen, Scott A. Mahlke, John C. Gyllenhaal, Wen-Mei W. Hwu |
ASPLOS | 2 |
| 1993 | Register Connection: A New Approach to Adding Registers into Instruction Set ArchitecturesabstractCode optimization and scheduling for superscalar and superpipelined processors often increase the register requirement of programs. For existing instruction sets with a small to moderate number of registers, this increased register requirement can be a factor that limits the effectivess of the compiler. In this paper, we introduce a new architectural method for adding a set of extended registers into an architecture. Using a novel concept of connection, this method allows the data stored in the extended registers to be accessed by instructions that apparently reference core registers. Furthermore, we address the technical issues involved in applying the new method to an architecture: instruction set extension, procedure call convention, context switching considerations, upward compatibility, efficient implementation, compiler support, and performance. Experimental results based on a prototype compiler and execution driven simulation show that the proposed method can significantly improve the performance of superscalar processors with a small or moderate number of registers. Tokuzo Kiyohara, Scott A. Mahlke, William Y. Chen, Roger A. Bringmann, Richard E. Hank, Sadun Anik, Wen-Mei W. Hwu |
ISCA | 3 |
| 1993 | The Effect of Code Expanding Optimizations on Instruction Cache DesignabstractShows that code expanding optimizations have strong and nonintuitive implications on instruction cache design. Three types of code expanding optimizations are studied in this paper: instruction placement, function inline expansion, and superscalar optimizations. Overall, instruction placement reduces the miss ratio of small caches. Function inline expansion improves the performance for small cache sizes, but degrades the performance of medium caches. Superscalar optimizations increase the miss ratio for all cache sizes. However, they also increase the sequentiality of instruction access so that a simple load forwarding scheme effectively cancels the negative effects. Overall, the authors show that with load forwarding, the three types of code expanding optimizations jointly improve the performance of small caches and have little effect on large caches.> William Y. Chen, Pohua P. Chang, Thomas M. Conte, Wen-Mei W. Hwu |
IEEE Trans. Computers | 1 |
| 1993 | The superblock: An effective technique for VLIW and superscalar compilation
Wen-Mei W. Hwu, Scott A. Mahlke, William Y. Chen, Pohua P. Chang, Nancy J. Warter, Roger A. Bringmann, Roland G. Ouellette, Richard E. Hank, Tokuzo Kiyohara, Grant E. Haab, John G. Holm, Daniel M. Lavery |
J. Supercomput. | 3 |
| 1993 | Sentinel Scheduling for VLIW and Superscalar ProcessorsabstractSpeculative execution is an important source of parallelism for VLIW and superscalar processors. A serious challenge with compiler-controlled speculative execution is to efficiently handle exceptions for speculative instructions. In this article, a set of architectural features and compile-time scheduling support collectively referred to assentinel schedulingis introduced. Sentinel scheduling provides an effective framework for both compiler-controlled speculative execution and exception handling. All program exceptions are accurately detected and reported in a timely manner with sentinel scheduling. Recovery from exceptions is also ensured with the model. Experimental results show the effectiveness of sentinel scheduling for exploiting instruction-level parallelism and overhead associated with exception handling. Scott A. Mahlke, William Y. Chen, Roger A. Bringmann, Richard E. Hank, Wen-Mei W. Hwu, Bob Rau, Mike Schlansker |
ACM Trans. Comput. Syst. | 2 |
| 1992 | Sentinel Scheduling for VLIW and Superscalar ProcessorsabstractSpeculative execution is an important source of parallelism for VLIW and superscalar processors. A serious challenge with compiler-controlled speculative execution is to accurately detect and report all program execution errors at the time of occurrence. In this paper, a set of architectural features and compile-time scheduling support referred to as sentinel scheduling is introduced. Sentinel scheduling provides an effective framework for compiler-controlled speculative execution that accurately detects and reports all exceptions. Sentinel scheduling also supports speculative execution of store instructions by providing a store buffer which allows probationary entries. Experimental results show that sentinel scheduling is highly effective for a wide range of VLIW and superscalar processors. Scott A. Mahlke, William Y. Chen, Wen-Mei W. Hwu, Bob Rau, Mike Schlansker |
ASPLOS | 2 |
| 1992 | Tolerating First Level Memory Access Latency in High-Performance Systems
William Y. Chen, Scott A. Mahlke, Wen-Mei W. Hwu |
ICPP (1) | 1 |
| 1992 | Tolerating data access latency with register preloadingabstractBy exploiting fine grain parallelism, superscalar processors can potentially increase the performance of future supercomputers. However, supercomputers typically have a long access delay to their first level memory which can severely restrict the performance of superscalar processors. Compilers attempt to move load instructions far enough ahead to hide this latency. However, conventional movement of load instructions is limited by data dependence analysis. This paper introduces a simple hardware scheme, referred to as preload register update, to allow the compiler to move load instructions even in the presence of inconclusive data dependence analysis results. Preload register update keeps the load destination registers coherent when load instructions are moved past store instructions that reference the same location. With this addition, superscalar processors can more effectively tolerate longer data access latencies. William Y. Chen, Scott A. Mahlke, Wen-Mei W. Hwu, Tokuzo Kiyohara, Pohua P. Chang |
ICS | 1 |
| 1992 | An efficient architecture for loop based data preloading
William Y. Chen, Roger A. Bringmann, Scott A. Mahlke, Richard E. Hank, James E. Sicolo |
MICRO | 1 |
| 1992 | Effective compiler support for predicated execution using the hyperblockabstractPredicated execution is an effective technique for dealing with conditional branches in application programs. However, there are several problems associated with conventional compiler support for predicated execution. First, all paths of control are combined into a single path regardless of their execution frequency and size with conventional if-conversion techniques. Second, speculative execution is difficult to combine with predicated execution. In this paper, we propose the use of a new structure, referred to as the hyperblock, to overcome these problems. The hyperblock is an efficient structure to utilize predicated execution for both compiletime optimization and scheduling. Preliminary experimental results show that the hyperblock is highly effective for a wide range of superscalar and VLIW processors. 1 Introduction Superscalar and VLIW processors can potentially provide large performance improvements over their scalar predecessors by providing multiple data paths and function u... Scott A. Mahlke, David C. Lin, William Y. Chen, Richard E. Hank, Roger A. Bringmann |
MICRO | 3 |
| 1992 | Compiler Code Transformations for Superscalar-Based High Performance SystemsabstractA set of compiler transformations designed to increase instruction-level parallelism is described. The effectiveness of these transformations is evaluated using 40 loop nests extracted from a range of supercomputer applications. This evaluation shows that increasing execution resources in superscalar/VLIW node processors yields little performance improvement unless loop unrolling and register renaming are applied. It also reveals that these two transformations are sufficient for DOALL loops. However, more advanced transformations are required in order for serial and DOACROSS loops to fully benefit from the increased execution resources. The results show that the six additional transformations studied satisfy most of this need.> Scott A. Mahlke, William Y. Chen, John C. Gyllenhaal, Wen-Mei W. Hwu |
SC | 2 |
| 1992 | Profile-guided Automatic Inline Expansion for C ProgramsabstractAbstract This paper describes critical implementation issues that must be addressed to develop a fully automatic inliner. These issues are: integration into a compiler, program representation, hazard prevention, expansion sequence control, and program modification. An automatic inter‐file inliner that uses profile information has been implemented and integrated into an optimizing C compiler. The experimental results show that this inliner achieves significant speedups for production C programs. Pohua P. Chang, Scott A. Mahlke, William Y. Chen, Wen-Mei W. Hwu |
Softw. Pract. Exp. | 3 |
| 1991 | The Effect of Compiler Optimizations on Available Parallelism in Scalar Programs
Scott A. Mahlke, Nancy J. Warter, William Y. Chen, Pohua P. Chang, Wen-Mei W. Hwu |
ICPP (2) | 3 |
| 1991 | IMPACT: An Architectural Framework for Multiple-Instruction-Issue ProcessorsabstractArticle Free Access Share on IMPACT: an architectural framework for multiple-instruction-issue processors Authors: Pohua P. Chang Center for Reliable and High-Performance Computing, University of Illinois, Urbana, IL Center for Reliable and High-Performance Computing, University of Illinois, Urbana, ILView Profile , Scott A. Mahlke Center for Reliable and High-Performance Computing, University of Illinois, Urbana, IL Center for Reliable and High-Performance Computing, University of Illinois, Urbana, ILView Profile , William Y. Chen Center for Reliable and High-Performance Computing, University of Illinois, Urbana, IL Center for Reliable and High-Performance Computing, University of Illinois, Urbana, ILView Profile , Nancy J. Warter Center for Reliable and High-Performance Computing, University of Illinois, Urbana, IL Center for Reliable and High-Performance Computing, University of Illinois, Urbana, ILView Profile , Wen-mei W. Hwu Center for Reliable and High-Performance Computing, University of Illinois, Urbana, IL Center for Reliable and High-Performance Computing, University of Illinois, Urbana, ILView Profile Authors Info & Claims ISCA '91: Proceedings of the 18th annual international symposium on Computer architectureApril 1991 Pages 266–275https://doi.org/10.1145/115952.115979Published:01 April 1991Publication History 215citation947DownloadsMetricsTotal Citations215Total Downloads947Last 12 Months59Last 6 weeks8 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Pohua P. Chang, Scott A. Mahlke, William Y. Chen, Nancy J. Warter, Wen-Mei W. Hwu |
ISCA | 3 |
| 1991 | Comparing Static and Dynamic Code Scheduling for Multiple-Instruction-Issue ProcessorsabstractThis paper examines two alternative approaches to supporting code scheduling for multiple-instruction-issue processors.One is to provide among the benchmark programs.To explain this variation, we have identified the conditions in these programs that make one approach perform better than the other. Pohua P. Chang, William Y. Chen, Scott A. Mahlke, Wen-Mei W. Hwu |
MICRO | 2 |
| 1991 | Data Access Microarchitectures for Superscalar Processors with Compiler-Assisted Data PrefetchingabstractThe performance of superscalar processors is more sensitive to the memory system delay than their single-issue predecessors. This paper examines alternative data access microarchitectures that effectively support compilerassisted data prefetching in superscalar processors. In particular, a prefetch buffer is shown to be more effective than increasing the cache dimension in solving the cache pollution problem. All in all, we show that a small data cache with compiler-assisted data prefetching can achieve a performance level close to that of an ideal cache. 1 Introduction Superscalar processors can potentially deliver more than five times speedup over conventional single-issue processors [1]. With the total execution cycle count dramatically reduced, each cycle becomes more significant to the overall performance. Because each data cache miss can introduce many extra execution cycles, a superscalar processor can easily lose the majority of its performance to the memory hierarchy. Out-of-... William Y. Chen, Scott A. Mahlke, Pohua P. Chang, Wen-Mei W. Hwu |
MICRO | 1 |