Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Trung A. Diep

dblp:40/2744 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
0since 2021 · last 2003
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Performance modeling and evaluation · 57% Processor architecture and microarchitecture · 27% Memory systems · 16%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Performance modeling and evaluation
benchmarking
0.012003
Scaling and Charact rizing Database Workloads: Bridging the Gap between Research and Practice · MICRO 2003
Performance modeling and evaluation
workload characterization
0.012003
Scaling and Charact rizing Database Workloads: Bridging the Gap between Research and Practice · MICRO 2003
Performance modeling and evaluation › simulation › processor simulation
microarchitecture simulation
0.011995
Performance Evaluation of the PowerPC 620 Microarchitecture · ISCA 1995
Processor architecture and microarchitecture
out-of-order execution
0.011995
Performance Evaluation of the PowerPC 620 Microarchitecture · ISCA 1995
Performance modeling and evaluation
simulation
0.011995
Performance Evaluation of the PowerPC 620 Microarchitecture · ISCA 1995
Processor architecture and microarchitecture › superscalar processor
superscalar pipeline
0.011995
Performance Evaluation of the PowerPC 620 Microarchitecture · ISCA 1995
Memory systems
cache
0.011993
EXPLORER: a retargetable and visualization-based trace-driven simulator for superscalar processors · MICRO 1993
Memory systems › cache design
data cache design
0.011993
EXPLORER: a retargetable and visualization-based trace-driven simulator for superscalar processors · MICRO 1993
Memory systems › cache › cache organization
multi-ported cache
0.011993
EXPLORER: a retargetable and visualization-based trace-driven simulator for superscalar processors · MICRO 1993
Processor architecture and microarchitecture
superscalar processor
0.011993
EXPLORER: a retargetable and visualization-based trace-driven simulator for superscalar processors · MICRO 1993
Processor architecture and microarchitecture
branch prediction
0.011995
Performance Evaluation of the PowerPC 620 Microarchitecture · ISCA 1995
Performance modeling and evaluation › simulation › discrete-event simulation
trace-driven simulation
0.011993
EXPLORER: a retargetable and visualization-based trace-driven simulator for superscalar processors · MICRO 1993

Methods — techniques the papers use, named apart from their topics

linear approximation · 0.0empirical measurement · 0.0trace-driven simulation · 0.0
YearPublicationVenuePosition
2003 Scaling and Charact rizing Database Workloads: Bridging the Gap between Research and Practice
abstract
On-line transaction processing (OLTP) workloads are crucial benchmarks for the design and analysis of server processors. Typical cached configurations used by researchers to simulate OLTP workloads are orders of magnitude smaller than the fully scaled configurations used by OEM vendors to achieve world-record transaction processing throughput. The objective of this study is to discover the underlying relationships that characterize OLTP performance over a wide range of configurations. To this end, we have derived the "iron law" of database performance. Using our iron law, we show that both the average instructions executed per transaction (IPX) and the average cycles per instruction (CPI) are critical to the transaction-throughput performance. We use an extensive, empirical examination of an Oracle based commercial OLTP workload on an Intel Xeon multiprocessor system to characterize the scaling behaviour of both the IPX and the CPI. We demonstrate that across a wide range of configurations the IPX and CPI behaviour follows predictable trends, which can be accurately characterized by simple linear or piece-wise linear approximations. Based on our data, we propose a method for selecting a minimal, representative workload configuration from which behaviours of much larger OLTP configurations can be accurately extrapolated.
Richard A. Hankins, Trung A. Diep, Murali Annavaram, Brian Hirano, Harald Eri, Hubert Nueckel, John Paul Shen
MICRO2
2002 Branch Behavior of a Commercial OLTP Workload on Intel IA32 Processors
abstract
This paper presents a detailed branch characterization of an Oracle based commercial on-line transaction processing workload, Oracle Database Benchmark (ODB), running on an IA32 processor. We ran a well-tuned ODB on Simics, a full system simulator, to collect the instruction traces used in this study. We compare the branch behavior of ODB with the branch behaviors of gcc, gzip and mcf from the SPECINT 2000 benchmark suite. Contrary to the popular belief that databases have unpredictable branches, we show that using larger predictors that capture enough branch history information, and using branch prediction schemes that reduce aliasing, conditional branches in ODB are more predictable than in gcc, gzip and mcf Due to frequent context switching in ODB, a hardware return address stack is ineffective in predicting return addresses for ODB. Based on further analysis, we propose and evaluate an enhanced return address predictor, which reduces return address mispredictions in ODB by 40%.
Murali Annavaram, Trung A. Diep, John Paul Shen
ICCD2
1995 Performance Evaluation of the PowerPC 620 Microarchitecture
abstract
The PowerPC 620™ microprocessor is the most recent and performance leading member of the PowerPC™ family. The 64-bit PowerPC 620 microprocessor employs a two-phase branch prediction scheme, dynamic renaming for all the register files, distributed multi-entry reservation stations, true out-of-order execution by six execution units, and a completion buffer for ensuring precise exceptions. This paper presents an instruction-level performance evaluation of the 620 microarchitecture. A performance simulator is developed using the VMW (Visualization-based Microarchitecture Workbench) retargetable framework. The VMW-based simulator accurately models the microarchitecture down to the machine cycle level. Extensive trace-driven simulation is performed using the SPEC92 benchmarks. Detailed quantitative analyses of the effectiveness of all key microarchitecture features are presented.
Trung A. Diep, Christopher Nelson, John Paul Shen
ISCA1
1993 Architecture-Compatible Code Boosting for Performance Enhancement of the IBM RS/6000
abstract
Boosting, first introduced by M.D. Smith et al. (1990), is an instruction scheduling technique that increases the instruction-level parallelism by allowing the compiler to move instructions speculatively up past conditional branches and providing hardware support to delay committing the side effects of the boosted instructions until the conditional branches have been resolved. The paper proposes an enhanced compilation technique similar to boosting that provides performance improvements while maintaining instruction set architecture compatibility and eliminating the need for complex hardware support. The technique, called architecture-compatible (AC) boosting, has been implemented for the IBM RS/6000 architecture. Code scheduling and machine simulation tools have been implemented, and experiments have been performed to demonstrate the feasibility of AC boosting on the current as well as future implementations of the IBM RS/6000 architecture.>
Trung A. Diep, Mikko H. Lipasti, John Paul Shen
ICCD1
1993 EXPLORER: a retargetable and visualization-based trace-driven simulator for superscalar processors
abstract
Superscalar implementations of RISC architectures are emerging as the dominant high-performance microprocessor technology for the mid-1990's. For instruction-level parallelism to increase beyond present levels, multiple memory operations per cycle are required. The paper evaluates several alternatives for two-ported data cache memory systems. A new split data cache memory design is compared to a more conventional true dual-ported memory. Experimental simulations are used to determine the performance benefits of these cache models on superscalar processors. These experiments are reported for a contemporary processor with modest instruction-level parallelism and for a hypothetical very aggressive, highly parallel processor.>
Trung A. Diep, John Paul Shen, Mike Phillip
MICRO1