Kevin M. Lepak

dblp:75/7031 · DBLP profile ↗
← Back
9ranked-venue papers
7as first author
0since 2021 · last 2005
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 6 first-authorSoftware engineering, systems software and programming languages · 3 · 3 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
6 papers
Memory systems · 59% Energy-efficient computing · 12% Processor architecture and microarchitecture · 10%

Topics — the 14 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
cache coherence
0.132002
Temporally silent stores · ASPLOS 2002
Silent Stores and Store Value Locality · IEEE Trans. Computers 2001
On the value locality of store instructions · ISCA 2000
Memory systems
silent store
0.132002
Temporally silent stores · ASPLOS 2002
Silent Stores and Store Value Locality · IEEE Trans. Computers 2001
On the value locality of store instructions · ISCA 2000
Memory systems › cache coherence
false sharing
0.122001
Silent Stores and Store Value Locality · IEEE Trans. Computers 2001
On the value locality of store instructions · ISCA 2000
Processor architecture and microarchitecture
value locality
0.122001
Silent Stores and Store Value Locality · IEEE Trans. Computers 2001
On the value locality of store instructions · ISCA 2000
Energy-efficient computing › power management
dynamic power management
0.112005
Temperature and supply Voltage aware performance and power modeling at microarchitecture level · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2005
Memory systems › cache coherence
coherence traffic reduction
0.012002
Temporally silent stores · ASPLOS 2002
Memory systems
memory access
0.012001
Silent Stores and Store Value Locality · IEEE Trans. Computers 2001
Electronic design automation
physical design
0.012001
Simultaneous Shield Insertion and Net Ordering under Explicit RLC Noise Constraint · DAC 2001
Memory systems › memory hierarchy
cache hierarchy
0.012000
Silent stores for free · MICRO 2000
Memory systems
memory access optimization
0.012000
Silent stores for free · MICRO 2000
Energy-efficient computing
voltage scaling
0.012005
Temperature and supply Voltage aware performance and power modeling at microarchitecture level · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2005
Parallel and multicore computing › parallel computing › parallel communication
multiprocessor communication
0.012002
Temporally silent stores · ASPLOS 2002
Hardware reliability and fault tolerance › error correction
cache error correction
0.012000
Silent stores for free · MICRO 2000
Hardware reliability and fault tolerance › error correction
error-correcting codes
0.012000
Silent stores for free · MICRO 2000

Methods — techniques the papers use, named apart from their topics

thermal modeling · 0.1leakage power modeling · 0.1MESI protocol extension · 0.0simulated annealing · 0.0producer-centric prediction · 0.0memory-centric prediction · 0.0simulation · 0.0
YearPublicationVenuePosition
2005 Reaping the Benefit of Temporal Silence to Improve Communication Performance
abstract
Communication misses - those serviced by dirty data in remote caches - are a pressing performance limiter in shared-memory multiprocessors. Recent research has indicated that temporally silent stores can be exploited to substantially reduce such misses, either with coherence protocol enhancements (MESTI); by employing speculation to create atomic silent store-pairs that achieve speculative lock elision (SLE); or by employing load value prediction (LVP). We evaluate all three approaches utilizing full-system, execution-driven simulation, with scientific and commercial workloads, to measure performance. Our studies indicate that accurate detection of elision idioms for SLE is vitally important for delivering robust performance and appears difficult for existing commercial codes. Furthermore, common datapath issues in out-of-order cores cause barriers to speculation and therefore may cause SLE failures unless SLE-specific speculation mechanisms are added to the microarchitecture. We also propose novel prediction and silence detection mechanisms that enable the MESTI protocol to deliver robust performance for all workloads. Finally, we conduct a detailed execution-driven performance evaluation of load value prediction (LVP), another simple method for capturing the benefit of temporally silent stores. We show that while theoretically LVP can capture the greatest fraction of communication misses among all approaches, it is usually not the most effective at delivering performance. This occurs because attempting to hide latency by speculating at the consumer, i.e. predicting load values, is fundamentally less effective than eliminating the latency at the source, by removing the invalidation effect of stores. Applying each method, we observe performance changes in application benchmarks ranging from 1% to 14% for an enhanced version of MESTI, -1.0% to 9% for LVP, -3% to 9% for enhanced SLE, and 2% to 21% for combined techniques
Kevin M. Lepak, Mikko H. Lipasti
ISPASS1
2005 Temperature and supply Voltage aware performance and power modeling at microarchitecture level
abstract
Performance and power are two primary design issues for systems ranging from server computers to handhelds. Performance is affected by both temperature and supply voltage because of the temperature and voltage dependence of circuit delay. Furthermore, as semiconductor technology scales down, leakage power's exponential dependence on temperature and supply voltage becomes significant. Therefore, future design studies call for temperature and voltage aware performance and power modeling. In this paper, we study microarchitecture-level temperature and voltage aware performance and power modeling. We present a leakage power model with temperature and voltage scaling, and show that leakage and total energy vary by 38% and 24%, respectively, between 65/spl deg/C and 110/spl deg/C. We study thermal runaway induced by the interdependence between temperature and leakage power, and demonstrate that without temperature-aware modeling, underestimation of leakage power may lead to the failure of thermal controls, and overestimation of leakage power may result in excessive performance penalties of up to 5.24%. All of these studies underscore the necessity of temperature-aware power modeling. Furthermore, we study optimal voltage scaling for best performance with dynamic power and thermal management under different packaging options. We show that dynamic power and thermal management allows designs to target at the common-case thermal scenario among benchmarks and improves performance by 6.59% compared to designs targeted at the worst case thermal scenario without dynamic power and thermal management. Additionally, the optimal V/sub dd/ for the best performance may not be the largest V/sub dd/ allowed by the given packaging platform, and that advanced cooling techniques can improve throughput significantly.
Weiping Liao, Lei He 0001, Kevin M. Lepak
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2004 Simultaneous shield insertion and net ordering for capacitive and inductive coupling minimization
abstract
In this article, we first show that existing net ordering formulations to minimize noise are no longer sufficient with the presence of inductive noise, and shield insertion is needed to minimize inductive noise. Using a K eff model as the figure of merit for inductive coupling, we then formulate two simultaneous shield insertion and net ordering (SINO) problems: the optimal SINO/NF problem to find a minimal area SINO solution that is free of capacitive and inductive noise, and the optimal SINO/NB problem to find a minimal area SINO solution that is free of capacitive noise and is under the given inductive noise bound. We reveal that both optimal SINO problems are NP-hard, and propose effective approximate algorithms for the two problems. Experiments show that our SINO/NB algorithm uses from 51% to 82% fewer shields compared to uniform shield insertion and net ordering (US + NO), and uses from 4% to 47% fewer shields compared to separated net ordering and shield insertion (NO + SI). Furthermore, the SINO/NB solutions under practical noise bounds use from 38% to 61% fewer shields compared to SINO/NF solutions, and use up to 36% fewer shields compared to the theoretical lower bound for optimal SINO/NF solutions. Moreover, we show that the K eff model has a high fidelity versus the noise voltage computed using accurate RLC circuit models and SPICE simulations. To the best of our knowledge, it is the first work that presents an in-depth study on the automatic layout optimization of multiple nets to minimize both capacitive and inductive noise.
Kevin M. Lepak, Jun Chen 0008, Lei He 0001
ACM Trans. Design Autom. Electr. Syst.1
2002 Temporally silent stores
abstract
Recent work has shown that silent stores--stores which write a value matching the one already stored at the memory location--occur quite frequently and can be exploited to reduce memory traffic and improve performance. This paper extends the definition of silent stores to encompass sets of stores that change the value stored at a memory location, but only temporarily, and subsequently return a previous value of interest to the memory location. The stores that cause the value to revert are called temporally silent stores. We redefine multiprocessor sharing to account for temporal silence and show that in the limit, up to 45% of communication misses in scientific and commercial applications can be eliminated by exploiting values that change only temporarily. We describe a practical mechanism that detects temporally silent stores and removes the coherence traffic they cause in conventional multiprocessors. We find that up to 42% of communication misses can be eliminated with a simple extension to the MESI protocol. Further, we examine application and operating system code to provide insight into the temporal silence phenomenon and characterize temporal silence by examining value frequencies and dynamic instruction distances between temporally silent pairs. These studies indicate that the operating system is involved heavily in temporal silence, in both commercial and scientific workloads, and that while detectable synchronization primitives provide substantial contributions, significant opportunity exists outside these references.
Kevin M. Lepak, Mikko H. Lipasti
ASPLOS1
2001 Simultaneous Shield Insertion and Net Ordering under Explicit RLC Noise Constraint
abstract
For multiple coupled RLC nets, we formulate the min-area simultaneous shield insertion and net ordering SINO/NB-ν problem to satisfy the given noise bound. We develop an efficient and conservative model to compute the peak noise, and apply the noise model to a simulated-annealing (SA) based algorithm for the SINO/NB-ν problem. Extensive and accurate experiments show that the SA-based algorithm is efficient, and always achieves solutions satisfying the given noise bound. It uses up to 71\% and 30\% fewer shields when compared to a greedy based shield insertion algorithm and a separated shield insertion and net ordering algorithm, respectively. To the best of our knowledge, it is the first work that presents an in-depth study on the min-area SINO problem under an explicit noise constraint.
Kevin M. Lepak, Irwan Luwandi, Lei He 0001
DAC1
2001 Silent Stores and Store Value Locality
abstract
Value locality, a recently discovered program attribute that describes the likelihood of the recurrence of previously seen program values, has been studied enthusiastically in the recent published literature. Much of the energy has focused on refining the initial efforts at predicting load instruction outcomes, with the balance of the effort examining the value locality of either all register-writing instructions or a focused subset of them. Surprisingly, there has been very little published characterization of or effort to exploit the value locality of data words stored to memory by computer programs. This paper presents such a characterization, including detailed source-level analysis of the causes of silent stores, proposes both memory-centric (based on message passing) and producer-centric (based on program structure) prediction mechanisms for stored data values, introduces the concept of silent stores and new definitions of multiprocessor false sharing based on these observations, and suggests new techniques for aligning cache coherence protocols and microarchitectural store handling techniques to exploit the value locality of stores. We find that realistic implementations of these techniques can significantly reduce multiprocessor data bus traffic and are more effective at reducing address bus traffic than the addition of Exclusive state to a MS I coherence protocol. We also show that squashing of silent stores can provide uniprocessor speedups greater than the addition of store-to-load forwarding.
Kevin M. Lepak, Gordon B. Bell, Mikko H. Lipasti
IEEE Trans. Computers1
2000 On the value locality of store instructions
abstract
Value locality, a recently discovered program attribute that describes the likelihood of the recurrence of previously-seen program values, has been studied enthusiastically in the recent published literature. Much of the energy has focused on refining the initial efforts at predicting load instruction outcomes, with the balance of the effort examining the value locality of either all register-writing instructions, or a focused subset of them. Surprisingly, there has been very little published characterization of or effort to exploit the value locality of data words stored to memory by computer programs. This paper presents such a characterization, proposes both memory-centric (based on message passing) and producer-centric (based on program structure) prediction mechanisms for stored data values, introduces the concept of silent stores and new definitions of multiprocessor false sharing based on these observations, and suggests new techniques for aligning cache coherence protocols and microarchitectural store handling techniques to exploit the value locality of stores. We find that realistic implementations of these techniques can significantly reduce multiprocessor data bus traffic and are more effective at reducing address bus traffic than the addition of Exclusive state to a MSI coherence protocol. We also show that squashing of silent stores can provide uniprocessor speedups greater than the addition of store-to-load forwarding.
Kevin M. Lepak, Mikko H. Lipasti
ISCA1
2000 Simultaneous shield insertion and net ordering for capacitive and inductive coupling minimization
abstract
In this paper, we first show that existing net ordering formulations to minimize noise are no longer valid with presence of inductive noise, and shield insertion is needed to minimize inductive noise.We then formulate two simultaneous shield insertion and net ordering (SINO) problems: the optimal SINO/NF problem to find a min-area SINO solution that is free of capacitive and inductive noise, and the optimal SINO/NB problem to find a min-area SINO solution that is free of capacitive noise and is under the given inductive noise bound.We reveal that both optimal SINO problems are NP-hard, and propose effective approximate algorithms for the two problems.Experiments show that our SINO/NB algorithm uses from 15% to 57% fewer shield wires when compared to separated net ordering and shield insertion procedure.Furthermore, under practical noise bounds, the SINO/NB solutions use from 44% to 67% fewer shield wires when compared to SINO/NF solutions, and use 10% to 40% fewer shield wires when compared to the theoretical lower bound for optimal SINO/NF solutions.Additionally, all our algorithms are extremely efficient to finish all examples in a few seconds.To the best of our knowledge, it is the first work that presents an indepth study on the simultaneous shield insertion and net ordering problem to minimize both capacitive and inductive noise.
Lei He 0001, Kevin M. Lepak
ISPD2
2000 Silent stores for free
abstract
Silent store instructions write values that exactly match the values that are already stored at the memory address that is being written. A recent study reveals that significant benefits can be gained by detecting and removing such stores from a program's execution. This paper studies the problem of detecting silent stores and shows that an average of 31% and 50% of silent stores can be detected for very low implementation cost, by exploiting temporal and spatial locality in a processor's load and store queues. We also show that over 83% of all silent stores can be detected using idle cache read access ports. Furthermore, we show that processors that use standard error-correction codes to protect data caches from transient errors can be modified only slightly to detect 100% of silent stores that hit in the cache. Finally, we show that silent store detection via these methods can result in a 11% harmonic mean performance improvement in a two-level store-through on-chip cache hierarchy that is based on a real microprocessor design.
Kevin M. Lepak, Mikko H. Lipasti
MICRO1