Pohua P. Chang

dblp:19/1698 · DBLP profile ↗
← Back
19ranked-venue papers
8as first author
0since 2021 · last 1995
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 16 · 6 first-authorSoftware engineering, systems software and programming languages · 7 · 3 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
10 papers
Processor architecture and microarchitecture · 82% Memory systems · 18%
Software engineering, system software, and programming languages
10 papers
Compilers and program optimization · 100%

Topics — the 22 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization
instruction scheduling
0.161995
Three Architecutral Models for Compiler-Controlled Speculative Execution · IEEE Trans. Computers 1995
The Importance of Prepass Code Scheduling for Superscalar and Superpipelined Processors · IEEE Trans. Computers 1995
Efficient Instruction Sequencing with Inline Target Insertion · IEEE Trans. Computers 1992
Processor architecture and microarchitecture
superscalar processor
0.031995
The Importance of Prepass Code Scheduling for Superscalar and Superpipelined Processors · IEEE Trans. Computers 1995
Data Access Microarchitectures for Superscalar Processors with Compiler-Assisted Data Prefetching · MICRO 1991
IMPACT: An Architectural Framework for Multiple-Instruction-Issue Processors · ISCA 1991
Processor architecture and microarchitecture
instruction-level parallelism
0.031995
Three Architecutral Models for Compiler-Controlled Speculative Execution · IEEE Trans. Computers 1995
Comparing Static and Dynamic Code Scheduling for Multiple-Instruction-Issue Processors · MICRO 1991
Exploiting Parallel Microprocessor Microarchitectures With a Compiler Code Generator · ISCA 1988
Compilers and program optimization › interprocedural optimization
inlining
0.021993
The Effect of Code Expanding Optimizations on Instruction Cache Design · IEEE Trans. Computers 1993
Inline Function Expansion for Compiling C Programs · PLDI 1989
Memory systems
cache
0.021991
Data Access Microarchitectures for Superscalar Processors with Compiler-Assisted Data Prefetching · MICRO 1991
Achieving High Instruction Cache Performance with an Optimizing Compiler · ISCA 1989
Compilers and program optimization › instruction scheduling
global instruction scheduling
0.011995
The Importance of Prepass Code Scheduling for Superscalar and Superpipelined Processors · IEEE Trans. Computers 1995
Compilers and program optimization › instruction scheduling
superblock scheduling
0.011995
Three Architecutral Models for Compiler-Controlled Speculative Execution · IEEE Trans. Computers 1995
Processor architecture and microarchitecture
instruction scheduling
0.011995
The Importance of Prepass Code Scheduling for Superscalar and Superpipelined Processors · IEEE Trans. Computers 1995
Processor architecture and microarchitecture
speculative execution
0.011995
Three Architecutral Models for Compiler-Controlled Speculative Execution · IEEE Trans. Computers 1995
Processor architecture and microarchitecture › pipelining
delayed branch
0.011992
Efficient Instruction Sequencing with Inline Target Insertion · IEEE Trans. Computers 1992
Memory systems › cache › CPU cache
instruction cache
0.021993
Achieving High Instruction Cache Performance with an Optimizing Compiler · ISCA 1989
The Effect of Code Expanding Optimizations on Instruction Cache Design · IEEE Trans. Computers 1993
Processor architecture and microarchitecture › pipelining
instruction pipeline
0.011992
Efficient Instruction Sequencing with Inline Target Insertion · IEEE Trans. Computers 1992
Processor architecture and microarchitecture › instruction-level parallelism › VLIW
VLIW processor
0.011991
IMPACT: An Architectural Framework for Multiple-Instruction-Issue Processors · ISCA 1991
Processor architecture and microarchitecture › superscalar processor
wide-issue processor
0.011991
IMPACT: An Architectural Framework for Multiple-Instruction-Issue Processors · ISCA 1991
Compilers and program optimization › compiler optimization
branch optimization
0.011989
Comparing Software and Hardware Schemes For Reducing the Cost of Branches · ISCA 1989
Compilers and program optimization
code layout optimization
0.011989
Achieving High Instruction Cache Performance with an Optimizing Compiler · ISCA 1989
Compilers and program optimization › code layout optimization
instruction placement
0.011989
Achieving High Instruction Cache Performance with an Optimizing Compiler · ISCA 1989
Processor architecture and microarchitecture
branch handling
0.011989
Comparing Software and Hardware Schemes For Reducing the Cost of Branches · ISCA 1989
Processor architecture and microarchitecture
pipelining
0.011989
Comparing Software and Hardware Schemes For Reducing the Cost of Branches · ISCA 1989
Memory systems
cache design
0.011993
The Effect of Code Expanding Optimizations on Instruction Cache Design · IEEE Trans. Computers 1993
Memory systems › cache management › cache interference
cache pollution
0.011991
Data Access Microarchitectures for Superscalar Processors with Compiler-Assisted Data Prefetching · MICRO 1991
Compilers and program optimization › dynamic optimization
profile-guided optimization
0.011989
Inline Function Expansion for Compiling C Programs · PLDI 1989

Methods — techniques the papers use, named apart from their topics

profile-based compilation · 0.0compiler-controlled speculation · 0.0performance evaluation · 0.0inline target insertion · 0.0benchmark comparison · 0.0simulation · 0.0profiling · 0.0optimizing compilation · 0.0code generation · 0.0compiler-assisted prefetching · 0.0profile information · 0.0
YearPublicationVenuePosition
1995 Using predicated execution to improve the performance of a dynamically scheduled machine with speculative execution
Po-Yung Chang, Eric Hao, Yale N. Patt, Pohua P. Chang
PACT4
1995 Profile-Guided Multi-Heuristic Branch Prediction
Pohua P. Chang, Utpal Banerjee
ICPP (1)1
1995 The Importance of Prepass Code Scheduling for Superscalar and Superpipelined Processors
abstract
Superscalar and superpipelined processors utilize parallelism to achieve peak performance that can be several times higher than that of conventional scalar processors. In order for this potential to be translated into the speedup of real program, the compiler must be able to schedule instructions so that the parallel hardware is effectively utilized. Previous work has shown that prepass code scheduling helps to produce a better schedule for scientific programs, but the importance of prescheduling has never been demonstrated for control-intensive non-numeric programs. These programs are significantly different from the scientific programs because they contain frequent branches. The compiler must do global scheduling in order to find enough independent instructions. In this paper, the code optimizer and scheduler of the IMPACT-I C compiler is described. Within this framework, we study the importance of prepass code scheduling for a set of production C programs. It is shown that, in contrast to the results previously obtained for scientific programs, prescheduling is not important for compiling control-intensive programs to the current generation of superscalar and superpipelined processors. However, if some of the current restrictions on upward code motion can be removed in future architectures, prescheduling would substantially improve the execution time of this class of programs on both superscalar and superpipelined processors.>
Pohua P. Chang, Daniel M. Lavery, Scott A. Mahlke, William Y. Chen, Wen-Mei W. Hwu
IEEE Trans. Computers1
1995 Three Architecutral Models for Compiler-Controlled Speculative Execution
abstract
To effectively exploit instruction level parallelism, the compiler must move instructions across branches. When an instruction is moved above a branch that it is control dependent on, it is considered to be speculatively executed since it is executed before it is known whether or not its result is needed. There are potential hazards when speculatively executing instructions. If these hazards can be eliminated, the compiler can more aggressively schedule the code. The hazards of speculative execution are outlined in this paper. Three architectural models: restricted, general, and boosting, which have increasing amounts of support for removing these hazards are discussed. The performance gained by each level of additional hardware support is analyzed using the IMPACT C compiler which performs superblock scheduling for superscalar and superpipelined processors.>
Pohua P. Chang, Nancy J. Warter, Scott A. Mahlke, William Y. Chen, Wen-Mei W. Hwu
IEEE Trans. Computers1
1993 The Effect of Code Expanding Optimizations on Instruction Cache Design
abstract
Shows that code expanding optimizations have strong and nonintuitive implications on instruction cache design. Three types of code expanding optimizations are studied in this paper: instruction placement, function inline expansion, and superscalar optimizations. Overall, instruction placement reduces the miss ratio of small caches. Function inline expansion improves the performance for small cache sizes, but degrades the performance of medium caches. Superscalar optimizations increase the miss ratio for all cache sizes. However, they also increase the sequentiality of instruction access so that a simple load forwarding scheme effectively cancels the negative effects. Overall, the authors show that with load forwarding, the three types of code expanding optimizations jointly improve the performance of small caches and have little effect on large caches.>
William Y. Chen, Pohua P. Chang, Thomas M. Conte, Wen-Mei W. Hwu
IEEE Trans. Computers2
1993 The superblock: An effective technique for VLIW and superscalar compilation
Wen-Mei W. Hwu, Scott A. Mahlke, William Y. Chen, Pohua P. Chang, Nancy J. Warter, Roger A. Bringmann, Roland G. Ouellette, Richard E. Hank, Tokuzo Kiyohara, Grant E. Haab, John G. Holm, Daniel M. Lavery
J. Supercomput.4
1992 Tolerating data access latency with register preloading
abstract
By exploiting fine grain parallelism, superscalar processors can potentially increase the performance of future supercomputers. However, supercomputers typically have a long access delay to their first level memory which can severely restrict the performance of superscalar processors. Compilers attempt to move load instructions far enough ahead to hide this latency. However, conventional movement of load instructions is limited by data dependence analysis. This paper introduces a simple hardware scheme, referred to as preload register update, to allow the compiler to move load instructions even in the presence of inconclusive data dependence analysis results. Preload register update keeps the load destination registers coherent when load instructions are moved past store instructions that reference the same location. With this addition, superscalar processors can more effectively tolerate longer data access latencies.
William Y. Chen, Scott A. Mahlke, Wen-Mei W. Hwu, Tokuzo Kiyohara, Pohua P. Chang
ICS5
1992 Profile-guided Automatic Inline Expansion for C Programs
abstract
Abstract This paper describes critical implementation issues that must be addressed to develop a fully automatic inliner. These issues are: integration into a compiler, program representation, hazard prevention, expansion sequence control, and program modification. An automatic inter‐file inliner that uses profile information has been implemented and integrated into an optimizing C compiler. The experimental results show that this inliner achieves significant speedups for production C programs.
Pohua P. Chang, Scott A. Mahlke, William Y. Chen, Wen-Mei W. Hwu
Softw. Pract. Exp.1
1992 Efficient Instruction Sequencing with Inline Target Insertion
abstract
Inline target insertion, a specific compiler and pipeline implementation method for delayed branches with squashing, is defined. The method is shown to offer two important features not discovered in previous studies. First, branches inserted into branch slots are correctly executed. Second, the execution returns correctly from interrupts or exceptions with only one program counter. These two features result in better performance and less software/hardware complexity than conventional delayed branching mechanisms.>
Wen-Mei W. Hwu, Pohua P. Chang
IEEE Trans. Computers2
1991 The Effect of Compiler Optimizations on Available Parallelism in Scalar Programs
Scott A. Mahlke, Nancy J. Warter, William Y. Chen, Pohua P. Chang, Wen-Mei W. Hwu
ICPP (2)4
1991 IMPACT: An Architectural Framework for Multiple-Instruction-Issue Processors
abstract
Article Free Access Share on IMPACT: an architectural framework for multiple-instruction-issue processors Authors: Pohua P. Chang Center for Reliable and High-Performance Computing, University of Illinois, Urbana, IL Center for Reliable and High-Performance Computing, University of Illinois, Urbana, ILView Profile , Scott A. Mahlke Center for Reliable and High-Performance Computing, University of Illinois, Urbana, IL Center for Reliable and High-Performance Computing, University of Illinois, Urbana, ILView Profile , William Y. Chen Center for Reliable and High-Performance Computing, University of Illinois, Urbana, IL Center for Reliable and High-Performance Computing, University of Illinois, Urbana, ILView Profile , Nancy J. Warter Center for Reliable and High-Performance Computing, University of Illinois, Urbana, IL Center for Reliable and High-Performance Computing, University of Illinois, Urbana, ILView Profile , Wen-mei W. Hwu Center for Reliable and High-Performance Computing, University of Illinois, Urbana, IL Center for Reliable and High-Performance Computing, University of Illinois, Urbana, ILView Profile Authors Info & Claims ISCA '91: Proceedings of the 18th annual international symposium on Computer architectureApril 1991 Pages 266–275https://doi.org/10.1145/115952.115979Published:01 April 1991Publication History 215citation947DownloadsMetricsTotal Citations215Total Downloads947Last 12 Months59Last 6 weeks8 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Pohua P. Chang, Scott A. Mahlke, William Y. Chen, Nancy J. Warter, Wen-Mei W. Hwu
ISCA1
1991 Comparing Static and Dynamic Code Scheduling for Multiple-Instruction-Issue Processors
abstract
This paper examines two alternative approaches to supporting code scheduling for multiple-instruction-issue processors.One is to provide among the benchmark programs.To explain this variation, we have identified the conditions in these programs that make one approach perform better than the other.
Pohua P. Chang, William Y. Chen, Scott A. Mahlke, Wen-Mei W. Hwu
MICRO1
1991 Data Access Microarchitectures for Superscalar Processors with Compiler-Assisted Data Prefetching
abstract
The performance of superscalar processors is more sensitive to the memory system delay than their single-issue predecessors. This paper examines alternative data access microarchitectures that effectively support compilerassisted data prefetching in superscalar processors. In particular, a prefetch buffer is shown to be more effective than increasing the cache dimension in solving the cache pollution problem. All in all, we show that a small data cache with compiler-assisted data prefetching can achieve a performance level close to that of an ideal cache. 1 Introduction Superscalar processors can potentially deliver more than five times speedup over conventional single-issue processors [1]. With the total execution cycle count dramatically reduced, each cycle becomes more significant to the overall performance. Because each data cache miss can introduce many extra execution cycles, a superscalar processor can easily lose the majority of its performance to the memory hierarchy. Out-of-...
William Y. Chen, Scott A. Mahlke, Pohua P. Chang, Wen-Mei W. Hwu
MICRO3
1991 Using Profile Information to Assist Classic Code Optimizations
abstract
Abstract This paper describes the design and implementation of an optimizing compiler that automatically generates profile information to assist classic code optimizations. This compiler contains two new components, an execution profiler and a profile‐based code optimizer, which are not commonly found in traditional optimizing compilers. The execution profiler inserts probes into the input program, executes the input program for several inputs, accumulates profile information and supplies this information to the optimizer. The profile‐based code optimizer uses the profile information to expose new optimization opportunities that are not visible to traditional global optimization methods. Experimental results show that the profile‐based code optimizer significantly improves the performance of production C programs that have already been optimized by a high‐quality global code optimizer.
Pohua P. Chang, Scott A. Mahlke, Wen-Mei W. Hwu
Softw. Pract. Exp.1
1989 Control flow optimization for supercomputer scalar processing
abstract
Control intensive scalar programs pose a very different challenge to highly pipelined supercomputers than vectorizable numeric applications. Function call/return and branch instructions disrupt the flow of instructions through the pipeline, degrading the utilization of the pipelined datapaths. This paper describes control flow optimization for scalar processing using an optimizing compiler. To obtain program control flow information, a system independent profiler has been integrated into the IMPACT-I C compiler. The control flow information obtained is converted into a weighted control graph. Based on the weighted control graph, function inline expansion, multi-way branch layout, and software branch prediction can be implemented. Using better compiler technology results in a very low cost hardware control unit (architecture) for high performance scalar processors.
Pohua P. Chang, Wen-Mei W. Hwu
ICS1
1989 Achieving High Instruction Cache Performance with an Optimizing Compiler
abstract
Increasing the execution power requires a high instruction issue bandwidth, and decreasing instruction encoding and applying some code improving techniques cause code expansion. Therefore, the instruction memory hierarchy performance has become an important factor of the system performance. An instruction placement algorithm has been implemented in the IMPACT-I (Illinois Microarchitecture Project using Advanced Compiler Technology - Stage I) C compiler to maximize the sequential and spatial localities, and to minimize mapping conflicts. This approach achieves low cache miss ratios and low memory traffic ratios for small, fast instruction caches with little hardware overhead. For ten realistic UNIX* programs, we report low miss ratios (average 0.5%) and low memory traffic ratios (average 8%) for a 2048-byte, direct-mapped instruction cache using 64-byte blocks. This result compares favorably with the fully associative cache results reported by other researchers. We also present the effect of cache size, block size, block sectoring, and partial loading on the cache performance. The code performance with instruction placement optimization is shown to be stable across architectures with different instruction encoding density.
Wen-Mei W. Hwu, Pohua P. Chang
ISCA2
1989 Comparing Software and Hardware Schemes For Reducing the Cost of Branches
abstract
Pipelining has become a common technique to increase throughput of the instruction fetch, instruction decode, and instruction execution portions of modern comput-ers. Branch instructions disrupt the flow of instructions through the the pipeline, increasing the overall execution cost of branch instructions. Three schemes to reduce the cost of branches are presented in the context of a gen-eral pipeline model. Ten realistic Unix domain programs are used to directly compare the cost and performance of the three schemes and the results are in favor of the software-based scheme. For example, the software-based scheme has a cost of 1.65 cycles/branch vs. a cost of 1.68 cycles/branch of the best hardware scheme for a highly pipelined processor (11-stage pipeline). The results are 1.19 (software scheme) vs. 1.23 cycles/branch (best hard-ware scheme) for a moderately pipelined processor (5-stage pipeline). 1
Wen-Mei W. Hwu, Thomas M. Conte, Pohua P. Chang
ISCA3
1989 Inline Function Expansion for Compiling C Programs
abstract
Inline function expansion replaces a function call with the function body. With automatic inline function expansion, programs can be constructed with many small functions to handle complexity and then rely on the compilation to eliminate most of the function calls. Therefore, inline expansion serves a tool for satisfying two conflicting goals: minizing the complexity of the program development and minimizing the function call overhead of program execution. A simple inline expansion procedure is presented which uses profile information to address three critical issues: code expansion, stack expansion, and unavailable function bodies. Experiments show that a large percentage of function calls/returns (about 59%) can be eliminated with a modest code expansion cost (about 17%) for twelve UNIX* programs.
Wen-Mei W. Hwu, Pohua P. Chang
PLDI2
1988 Exploiting Parallel Microprocessor Microarchitectures With a Compiler Code Generator
abstract
Several experiments using a versatile optimizing compiler to evaluate the benefit of four forms of microarchitectural parallelisms (multiple microoperations issued per cycle, multiple result-distribution buses, multiple execution units, and pipelined execution units) are described. The first 14 Livermore loops and 10 of the linpack subroutines are used as the preliminary benchmarks. The compiler generates optimized code for different microarchitecture configurations. It is shown how the compiler can help to derive a balanced design for high performance. For each given set of technology constraints, these experiments can be used to derive a cost-effective microarchitecture to execute each given set of workload programs at high speed.>
Wen-Mei W. Hwu, Pohua P. Chang
ISCA2