EDBT 2026 Demo / reviewers in the wild / expert
Shih-Lien Lu
dblp:35/5292
· DBLP profile ↗
51ranked-venue papers
6as first author
0since 2021 · last 2016
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 47 · 6 first-authorSoftware engineering, systems software and programming languages · 7Applied, interdisciplinary, general and emerging computing · 3Computer networks · 2Security and privacy · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
24 papers |
Memory systems · 41% Hardware reliability and fault tolerance · 15% Energy-efficient computing · 15% | |
| Computer networks
1 paper |
Routing and switching · 100% |
Topics — the 30 heaviest of 57, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems
DRAM |
0.9 | 4 | 2016 | DRAM Refresh Mechanisms, Penalties, and Trade-Offs · IEEE Trans. Computers 2016 Improving DRAM latency with dynamic asymmetric subarray · MICRO 2015 Flexible auto-refresh: enabling scalable and energy-efficient DRAM refresh reductions · ISCA 2015 |
Hardware reliability and fault tolerance › memory reliability
cache reliability |
0.3 | 3 | 2011 | Energy-efficient cache design using variable-strength error-correcting codes · ISCA 2011 Improving cache lifetime reliability at ultra-low voltages · MICRO 2009 Trading off Cache Capacity for Reliability to Enable Low Voltage Operation · ISCA 2008 |
Hardware reliability and fault tolerance › error correction
error-correcting codes |
0.3 | 3 | 2011 | Energy-efficient cache design using variable-strength error-correcting codes · ISCA 2011 Reducing cache power with low-cost, multi-bit error-correcting codes · ISCA 2010 Improving cache lifetime reliability at ultra-low voltages · MICRO 2009 |
Electronic design automation
high-level synthesis |
0.2 | 2 | 2011 | Automatic Pipelining From Transactional Datapath Specifications · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2011 Automatic multithreaded pipeline synthesis from transactional datapath specifications · DAC 2010 |
Memory systems
access latency reduction |
0.2 | 1 | 2015 | Improving DRAM latency with dynamic asymmetric subarray · MICRO 2015 |
Memory systems › DRAM
DRAM refresh |
0.2 | 1 | 2015 | Flexible auto-refresh: enabling scalable and energy-efficient DRAM refresh reductions · ISCA 2015 |
Memory systems
cache |
0.2 | 2 | 2014 | Sandbox Prefetching: Safe run-time evaluation of aggressive prefetchers · HPCA 2014 Design, implementation, and verification of active cache emulator (ACE) · FPGA 2006 |
Memory systems › DRAM › DRAM refresh
refresh energy reduction |
0.2 | 2 | 2013 | Technology comparison for large last-level caches (L3Cs): Low-leakage SRAM, low write-energy STT-RAM, and refresh-optimized eDRAM · HPCA 2013 Reducing cache power with low-cost, multi-bit error-correcting codes · ISCA 2010 |
Memory systems › cache › prefetching
hardware prefetching |
0.2 | 1 | 2014 | Sandbox Prefetching: Safe run-time evaluation of aggressive prefetchers · HPCA 2014 |
Processor architecture and microarchitecture
latency hiding |
0.2 | 1 | 2014 | Sandbox Prefetching: Safe run-time evaluation of aggressive prefetchers · HPCA 2014 |
Energy-efficient computing
voltage scaling |
0.2 | 3 | 2011 | Energy-efficient cache design using variable-strength error-correcting codes · ISCA 2011 Adaptive Cache Design to Enable Reliable Low-Voltage Operation · IEEE Trans. Computers 2011 Trading off Cache Capacity for Reliability to Enable Low Voltage Operation · ISCA 2008 |
Routing and switching › IP lookup
hash-based lookup |
0.2 | 1 | 2013 | Guided multiple hashing: Achieving near perfect balance for fast routing lookup · ICNP 2013 |
Routing and switching
IP lookup |
0.2 | 1 | 2013 | Guided multiple hashing: Achieving near perfect balance for fast routing lookup · ICNP 2013 |
Memory systems › DRAM › DRAM architecture
embedded DRAM |
0.2 | 1 | 2013 | Technology comparison for large last-level caches (L3Cs): Low-leakage SRAM, low write-energy STT-RAM, and refresh-optimized eDRAM · HPCA 2013 |
Memory systems › memory hierarchy › cache hierarchy
last-level cache |
0.2 | 1 | 2013 | Technology comparison for large last-level caches (L3Cs): Low-leakage SRAM, low write-energy STT-RAM, and refresh-optimized eDRAM · HPCA 2013 |
Processor architecture and microarchitecture › pipelining
pipeline design |
0.2 | 2 | 2011 | Automatic Pipelining From Transactional Datapath Specifications · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2011 Automatic multithreaded pipeline synthesis from transactional datapath specifications · DAC 2010 |
Processor architecture and microarchitecture
pipelining |
0.1 | 2 | 2011 | Automatic Pipelining From Transactional Datapath Specifications · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2011 Advances of the Counterflow Pipeline Microarchitecture · HPCA 1997 |
Memory systems
cache design |
0.1 | 1 | 2011 | Adaptive Cache Design to Enable Reliable Low-Voltage Operation · IEEE Trans. Computers 2011 |
Hardware reliability and fault tolerance
error correction |
0.1 | 1 | 2011 | Adaptive Cache Design to Enable Reliable Low-Voltage Operation · IEEE Trans. Computers 2011 |
Energy-efficient computing
low-voltage operation |
0.1 | 2 | 2011 | Trading off Cache Capacity for Reliability to Enable Low Voltage Operation · ISCA 2008 Adaptive Cache Design to Enable Reliable Low-Voltage Operation · IEEE Trans. Computers 2011 |
Energy-efficient computing › power management › memory power management
cache energy reduction |
0.1 | 1 | 2010 | Reducing cache power with low-cost, multi-bit error-correcting codes · ISCA 2010 |
Memory systems › DRAM › DRAM refresh
eDRAM refresh |
0.1 | 1 | 2010 | Reducing cache power with low-cost, multi-bit error-correcting codes · ISCA 2010 |
Hardware reliability and fault tolerance › error correction
multi-bit error correction |
0.1 | 1 | 2010 | Reducing cache power with low-cost, multi-bit error-correcting codes · ISCA 2010 |
Electronic design automation › high-level synthesis
pipeline synthesis |
0.1 | 1 | 2010 | Automatic multithreaded pipeline synthesis from transactional datapath specifications · DAC 2010 |
Energy-efficient computing › power management
dynamic voltage and frequency scaling |
0.1 | 1 | 2009 | Circuit techniques for dynamic variation tolerance · DAC 2009 |
Integrated circuit design › variation-aware design
variation-tolerant circuit design |
0.1 | 1 | 2009 | Circuit techniques for dynamic variation tolerance · DAC 2009 |
Hardware reliability and fault tolerance
soft errors |
0.1 | 3 | 2011 | Energy-efficient cache design using variable-strength error-correcting codes · ISCA 2011 Improving cache lifetime reliability at ultra-low voltages · MICRO 2009 Coming challenges in microarchitecture and architecture · Proc. IEEE 2001 |
Energy-efficient computing › power management › memory power management
DRAM power reduction |
0.1 | 1 | 2016 | DRAM Refresh Mechanisms, Penalties, and Trade-Offs · IEEE Trans. Computers 2016 |
Energy-efficient computing
power management |
0.1 | 1 | 2016 | DRAM Refresh Mechanisms, Penalties, and Trade-Offs · IEEE Trans. Computers 2016 |
Performance modeling and evaluation › simulation
architectural simulation |
0.1 | 1 | 2007 | An FPGA-based Pentium in a complete desktop system · FPGA 2007 |
Methods — techniques the papers use, named apart from their topics
simulation · 0.4mechanism categorization · 0.2experimental characterization · 0.2circuit simulation · 0.2dynamic asymmetric subarray · 0.2offset prefetching · 0.2component-based modeling · 0.2bloom filter · 0.2technology modeling · 0.2full-system simulation · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2016 | Small cache lookaside table for fast DRAM cache accessabstractLarge off-die stacked DRAM caches have been proposed to provide higher effective bandwidth and lower average latency to main memory. Designing a large off-die DRAM cache with conventional block size requires a large tag array which is impractical to fit on-die. Placing the large directory off-die prolong the latency since a tag access is necessary before the data can be accessed. This additional trip also generates extra off-die traffic. In this paper, we present a novel design called Cache Lookaside Table (CLT) to reduce the average access latency and to lessen off-die tag array accesses. The basic approach is to cache a small amount of recently referenced tags on-die. An off-die tag access is avoided when a requested block's tag hits a cached tag. To save on-die space, cached tags are recorded in a large sector for sharing tags with multiple blocks. However, due to the loss of one-to-one physical mapping of the cached tags and the data array, a way pointer is added for each block to indicate its way location. The proposed CLT exploits memory reference locality and provides a fast alternative tag path to capture most of the DRAM cache requests. In comparison with other proposed DRAM caching mechanisms, the on-die CLT approach shows an average performance improvement in the range of 4-15%. Qi Zeng 0006, Jih-Kwon Peir, Shih-Lien Lu |
IPCCC | 4 |
| 2016 | Runahead Cache Misses Using Bloom FilterabstractIn order to hide long memory latency and alleviate memory bandwidth requirement, a fourth-level cache (L4) is introduced in modern high-performance multi-core systems for supporting parallel computation. However, additional cache level causes higher cache miss penalty since a request needs to go through all levels of caches to reach to the main memory. In this paper, we introduce a new way of using a Bloom Filter (BF) to predict cache misses at any cache level in a multicore system. These misses can runahead to access lower-level caches or memory to reduce the miss penalty. The proposed hashing scheme extends the cache index of the target set and uses it for accessing the BF array to avoid counters in the BF array. Performance evaluation using a set of SPEC2006 benchmarks on 8-core systems with 4-level cache hierarchy shows that using a BF for the third-level (L3) cache to filter and runahead L3 misses, the IPCs can be improved by 4-20% with an average 10.5%. In comparison with the delay-recalibration scheme, the improvement is 3.5-4.8%. Qi Zeng 0006, Jih-Kwon Peir, Shih-Lien Lu |
PDCAT | 4 |
| 2016 | DRAM Refresh Mechanisms, Penalties, and Trade-OffsabstractEver-growing application data footprints demand faster main memory with larger capacity. DRAM has been the technology choice for main memory due to its low latency and high density. However, DRAM cells must be refreshed periodically to preserve their content. Refresh operations negatively affect performance and power. Traditionally, the performance and power overhead of refresh have been insignificant. But as the size and speed of DRAM chips continue to increase, refresh becomes a dominating factor of DRAM performance and power dissipation. In this paper, we conduct a comprehensive study of the issues related to refresh operations in modern DRAMs. Specifically, we describe the difference in refresh operations between modern synchronous DRAM and traditional asynchronous DRAM; the refresh modes and timings; and variations in data retention time. Moreover, we quantify refresh penalties versus device speed, size, and total memory capacity. We also categorize refresh mechanisms based on command granularity, and summarize refresh techniques proposed in research papers. Finally, based on our experiments and observations, we propose guidelines for mitigating DRAM refresh penalties. Ishwar Bhati, Mu-Tien Chang, Zeshan Chishti, Shih-Lien Lu, Bruce L. Jacob |
IEEE Trans. Computers | 4 |
| 2015 | Flexible auto-refresh: enabling scalable and energy-efficient DRAM refresh reductionsabstractDRAM cells require periodic refreshing to preserve data. In JEDEC DDRx devices, a refresh operation is performed via an auto-refresh command, which refreshes multiple rows in multiple banks simultaneously. The internal implementation of auto-refresh is completely opaque outside the DRAM --- all the memory controller can do is to instruct the DRAM to refresh itself --- the DRAM handles all else, in particular determining which rows in which banks are to be refreshed. This is in conflict with a large body of research on reducing the refresh overhead, in which the memory controller needs fine-grained control over which regions of the memory are refreshed. For example, prior works exploit the fact that a subset of DRAM rows can be refreshed at a slower rate than other rows due to access rate or retention period variations. However, such row-granularity approaches cannot use the standard auto-refresh command, which refreshes an entire batch of rows at once and does not permit skipping of rows. Consequently, prior schemes are forced to use explicit sequences of activate (ACT) and precharge (PRE) operations to mimic row-level refreshing. The drawback is that, compared to using JEDEC's auto-refresh mechanism, using explicit ACT and PRE commands is inefficient, both in terms of performance and power. Ishwar Bhati, Zeshan Chishti, Shih-Lien Lu, Bruce L. Jacob |
ISCA | 3 |
| 2015 | Improving DRAM latency with dynamic asymmetric subarrayabstractThe evolution of DRAM technology has been driven by capacity and bandwidth during the last decade. In contrast, DRAM access latency stays relatively constant and is trending to increase. Much efforts have been devoted to tolerate memory access latency but these techniques have reached the point of diminishing returns. Having shorter bitline and wordline length in a DRAM device will reduce the access latency. However by doing so it will impact the array efficiency. In the mainstream market, manufacturers are not willing to trade capacity for latency. Prior works had proposed hybrid-bitline DRAM design to overcome this problem. However, those methods are either intrusive to the circuit and layout of the DRAM design, or there is no direct way to migrate data between the fast and slow levels. Shih-Lien Lu, Ying-Chen Lin, Chia-Lin Yang |
MICRO | 1 |
| 2014 | Sandbox Prefetching: Safe run-time evaluation of aggressive prefetchersabstractMemory latency is a major factor in limiting CPU performance, and prefetching is a well-known method for hiding memory latency. Overly aggressive prefetching can waste scarce resources such as memory bandwidth and cache capacity, limiting or even hurting performance. It is therefore important to employ prefetching mechanisms that use these resources prudently, while still prefetching required data in a timely manner. In this work, we propose a new mechanism to determine at run-time the appropriate prefetching mechanism for the currently executing program, called Sandbox Prefetching. Sandbox Prefetching evaluates simple, aggressive offset prefetchers at run-time by adding the prefetch address to a Bloom filter, rather than actually fetching the data into the cache. Subsequent cache accesses are tested against the contents of the Bloom filter to see if the aggressive prefetcher under evaluation could have accurately prefetched the data, while simultaneously testing for the existence of prefetchable streams. Real prefetches are performed when the accuracy of evaluated prefetchers exceeds a threshold. This method combines the ideas of global pattern confirmation and immediate prefetching action to achieve high performance. Sandbox Prefetching improves performance across the tested workloads by 47.6% compared to not using any prefetching, and by 18.7% compared to the Feedback Directed Prefetching technique. Performance is also improved by 1.4% compared to the Access Map Pattern Matching Prefetcher, while incurring considerably less logic and storage overheads. Seth H. Pugsley, Zeshan Chishti, Chris Wilkerson, Peng-fei Chuang, Robert L. Scott, Aamer Jaleel, Shih-Lien Lu, Kingsum Chow, Rajeev Balasubramonian |
HPCA | 7 |
| 2014 | Recycled Error Bits: Energy-Efficient Architectural Support for Floating Point AccuracyabstractIn this work, we provide energy-efficient architectural support for floating point accuracy. For each floating point addition performed, we "recycle" that operation's rounding error. We make this error architecturally visible such that it can be used, whenever desired, by software. We also design a compiler pass that allows software to automatically use this feature. Experimental results on physical hardware show that software that exploits architecturally recycled error bits can (a) achieve accuracy comparable to a 64-bit FPU with performance and energy that are comparable to a 32-bit FPU, and (b) achieve accuracy comparable to an all-software scheme for 128-bit accuracy with far better performance and energy usage. Ralph Nathan, Bryan Anthonio, Shih-Lien Lu, Helia Naeimi, Daniel J. Sorin, Xiaobai Sun |
SC | 3 |
| 2014 | DArT: A Component-Based DRAM Area, Power, and Timing Modeling ToolabstractDRAM renovation calls for a holistic architecture exploration to cope with bandwidth growth and latency reduction need. In this paper, we present DRAM area power timing (DArT), a DRAM area, power, and timing modeling tool, for array assembly and interface customization. Through proper design abstraction, our component-based modeling approach provides increased flexibility and higher accuracy, making DArT suitable for DRAM architecture exploration and performance estimation. We validate the accuracy of DArT with respect to the physical layout and circuit simulation of an industrial 68 nm commodity DRAM device as a reference. The experiment results show that the maximum deviations from the reference design, in terms of area, timing, and power, are 3.2%, 4.92%, and 1.73%, respectively. For an architectural projection by porting it to a 45 nm process, the maximum deviations are 3.4%, 3.42%, and 8.57%, respectively. The combination of modeling performance, flexibility, and accuracy of DArT allows us to easily explore new DRAM architectures in the future, including 3-D stacked DRAM. Hsiu-Chuan Shih, Pei-Wen Luo, Jen-Chieh Yeh, Shu-Yen Lin, Ding-Ming Kwai, Shih-Lien Lu, Andre Schaefer, Cheng-Wen Wu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2013 | Technology comparison for large last-level caches (L3Cs): Low-leakage SRAM, low write-energy STT-RAM, and refresh-optimized eDRAMabstractLarge last-level caches (L3Cs) are frequently used to bridge the performance and power gap between processor and memory. Although traditional processors implement caches as SRAMs, technologies such as STT-RAM (MRAM), and eDRAM have been used and/or considered for the implementation of L3Cs. Each of these technologies has inherent weaknesses: SRAM is relatively low density and has high leakage current; STT-RAM has high write latency and write energy consumption; and eDRAM requires refresh operations. As future processors are expected to have larger last-level caches, the goal of this paper is to study the trade-offs associated with using each of these technologies to implement L3Cs. In order to make useful comparisons between SRAM, STTRAM, and eDRAM L3Cs, we model them in detail and apply low power techniques to each of these technologies to address their respective weaknesses. We optimize SRAM for low leakage and optimize STT-RAM for low write energy. Moreover, we classify eDRAM refresh-reduction schemes into two categories and demonstrate the effectiveness of using dead-line prediction to eliminate unnecessary refreshes. A comparison of these technologies through full-system simulation shows that the proposed refresh-reduction method makes eDRAM a viable, energy-efficient technology for implementing L3Cs. Mu-Tien Chang, Paul Rosenfeld, Shih-Lien Lu, Bruce L. Jacob |
HPCA | 3 |
| 2013 | Guided multiple hashing: Achieving near perfect balance for fast routing lookupabstractThe routing and packet forwarding function is at the core of the IP network-layer protocols. The throughput of a router is constrained by the speed at which the routing table lookup can be performed. Hash-based lookup has been a research focus in this area due to its O(1) average lookup time, as compared to other approachs such as trie-based lookup which tends to make more memory accesses. With a series of prior multi-hashing developments, including d-random, 2-left, and d-left, we discover that a new guided multi-hashing approach holds the promise of further pushing the envelope of this line of research to make significant performance improvement beyond what today's best technology can achieve. Our guided multi-hashing approach achieves near perfect load balance among hash buckets, while limiting the number of buckets to be probed for each key (address) lookup, where each bucket holds one or a few routing entries. Unlike the localized optimization by the prior approaches, we utilize the full information of multi-hash mapping from keys to hash buckets for global key-to-bucket assignment. We have dual objectives of lowering the bucket size while increasing empty buckets, which helps to reduce the number of buckets brought from off-chip memory to the network processor for each lookup. We introduce mechanisms to make sure that most lookups only require one bucket to be fetched. Our simulation results show that with the same number of hash functions, the guided multiple-hashing schemes are more balanced than d-left and others, while the average number of buckets to be accessed for each lookup is reduced by 20–50%. Jih-Kwon Peir, Shigang Chen, Shih-Lien Lu |
ICNP | 6 |
| 2013 | Guided Region-Based GPU Scheduling: Utilizing Multi-thread Parallelism to Hide Memory LatencyabstractModern General-Purpose computation on Graphics Processing Units (GPGPUs) explore parallelism in applications by building massively parallel architecture and apply multithreading technology to hide the instruction and memory latencies. Such architectures become increasingly popular for parallel applications using CUDA/OpenCL programming languages. In this paper, we investigate thread scheduling algorithms on such highly-threaded GPGPUs. The traditional round-robin scheduling schemes are inefficient in handling instruction execution and memory accesses with disparate latencies. We introduce a new GPGPU thread (warp) scheduling algorithm which enables flexible roundrobin distance for efficiently utilizing multithread parallelism and use program-guided priority shift among concurrent threads (warps) to allow more overlaps between short-latency compute instructions and long-latency memory accesses. Performance evaluations demonstrate that the new scheduling algorithm improves a set of kernel execution times by an average of 12% with 52% reduction on scheduler stall cycles over the fine-granularity round-robin scheme. In this paper, we also accomplish a thorough evaluation of various thread scheduling algorithms based on the amount of hardware threads, the scheduling overhead, and the global memory latency. Jianmin Chen, Jih-Kwon Peir, Shih-Lien Lu |
IPDPS | 6 |
| 2013 | Reducing cache and TLB power by exploiting memory region and privilege level semantics
Zhen Fang 0002, Li Zhao 0002, Xiaowei Jiang, Shih-Lien Lu, Ravi R. Iyer 0001, Tong Li 0003 |
J. Syst. Archit. | 4 |
| 2012 | Design for test and reliability in ultimate CMOSabstractThis session brings together specialists from the DfT, DfY and DfR domains that will address key problems together with their solutions for the 14 nm node and beyond, dealing with extremely complex chips affected by high defect levels, unpredictable and heterogeneous timing behavior, circuit degradation over time, including extreme situations related with the ultimate CMOS nodes, where all processor nodes, routers and links of single-chip massively parallel tera-device processors could comprise timing faults (such as delay faults or clock skews); a large percentage of these parts are affected by catastrophic failures; all parts experience significant performance degradations over time; and new catastrophic failures occur at low MTBF. Michael Nicolaidis, Lorena Anghel, Nacer-Eddine Zergainoh, Yervant Zorian, Tanay Karnik, Keith A. Bowman, James W. Tschanz, Shih-Lien Lu, Carlos Tokunaga, Arijit Raychowdhury, Muhammad M. Khellah, Jaydeep P. Kulkarni, Vivek De, Dimiter R. Avresky |
DATE | 8 |
| 2012 | Scaling the "Memory Wall": Designer trackabstractDRAM has been the technology for computer main memory since Intel released the first commercial DRAM chip (i1103) in 1970. As technology scales and demand for memory performance, it seems DRAM is facing several challenges. Many other memory technologies are anticipated to replace it but none has emerged as a clear winner thus far. In this paper we post the question. Is it possible to re-examine the design of DRAM to continue its life for another decade at least? Shih-Lien Lu, Tanay Karnik, Ganapati Srinivasa, Kai-Yuan Chao, Doug Carmean, Jim Held |
ICCAD | 1 |
| 2012 | Reducing L1 caches power by exploiting software semanticsabstractTo access a set-associative L1 cache in a high-performance processor, all ways of the selected set are searched and fetched in parallel using physical address bits. Such a cache is oblivious of memory references' software semantics such as stack-heap bifurcation of the memory space, and user-kernel ring levels. This constitutes a waste of energy since e.g., a user-mode instruction fetch will never hit a cache block that contains kernel code. Similarly, a stack access will not hit a cacheline that contains heap data. Zhen Fang 0002, Li Zhao 0002, Xiaowei Jiang, Shih-Lien Lu, Ravi R. Iyer 0001, Tong Li 0003 |
ISLPED | 4 |
| 2012 | Direct Compare of Information Coded With Error-Correcting CodesabstractThere are situations in a computing system where incoming information needs to be compared with a piece of stored data to locate the matching entry, e.g., cache tag array lookup and translation look-aside buffer matching. If the stored data is protected with error-correcting codes (ECC) for reliability reason, the previous solution is to access the stored information, decode and correct if necessary before it is used to compare with the incoming data. The decoding and correcting step increases the total access time, which is often critical. In this paper, we propose a method to improve the compare latency for information encoded with ECC. We use the cache tag array look-up as an example, and results show that 30% gate count reduction and 12% latency reduction are achieved. Wei Wu 0024, Dinesh Somasekhar, Shih-Lien Lu |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2011 | Distributed hardware matcher framework for SoC survivabilityabstractModern systems on chip (SoCs) are rapidly becoming complex high-performance computational devices, featuring multiple general purpose processor cores and a variety of functional IP blocks, communicating with each other through on-die fabric. While modular SoC design provides power savings and simplifies the development process, it also leaves significant room for a special type of hardware bugs, interaction errors, to slip through pre- and post-silicon verification. Consequently, hard to fix silicon escapes may be discovered late in production schedule or even after a market release, potentially causing costly delays or recalls. In this work we propose a unified error detection and recovery framework that incorporates programmable features into the on-die fabric of an SoC, so triggers of escaped interaction bugs can be detected at runtime. Furthermore, upon detection, our solution locks the interface of an IP for a programmed time period, thus altering interactions between accesses and bypassing the bug in a manner transparent to software. For classes of errors that cannot be circumvented by this in-hardware technique our framework is programmed to propagate the error detection to the software layer. Our experiments demonstrate that the proposed framework is capable of detecting a range of interaction errors with less than 0.01% performance penalty and 0.45% area overhead. Ilya Wagner, Shih-Lien Lu |
DATE | 2 |
| 2011 | Energy-efficient cache design using variable-strength error-correcting codesabstractVoltage scaling is one of the most effective mechanisms to improve microprocessors' energy efficiency. However, processors cannot operate reliably below a minimum voltage, Vccmin, since hardware structures may fail. Cell failures in large memory arrays (e.g., caches) typically determine Vccmin for the whole processor. We observe that most cache lines exhibit zero or one failures at low voltages. However, a few lines, especially in large caches, exhibit multi-bit failures and increase Vccmin. Previous solutions either significantly reduce cache capacity to enable uniform error correction across all lines, or significantly increase latency and bandwidth overheads when amortizing the cost of error-correcting codes (ECC) over large lines. Alaa R. Alameldeen, Ilya Wagner, Zeshan Chishti, Wei Wu 0024, Chris Wilkerson, Shih-Lien Lu |
ISCA | 6 |
| 2011 | Adaptive Cache Design to Enable Reliable Low-Voltage OperationabstractThe performance/energy trade-off is widely acknowledged as a primary design consideration for modern processors. A less discussed, though equally important, trade-off is the reliability/energy trade-off. Many design features that increase reliability (e.g., redundancy, error detection, and correction) have the side effect of consuming more energy. Many energy-saving features (e.g., voltage scaling) have the side effect of making systems less reliable. In this paper, we propose an adaptive cache design that enables the operating system to optimize for performance or energy efficiency without sacrificing reliability. Our proposed mechanism enables a cache with a wide operating range, where the cache can use a variable part of its data array to store error-correcting codes. A reliable, energy-efficient cache can use up to half of its data array to store error-correcting codes so that it can reliably operate at a low voltage to reduce energy. A reliable high-performance cache uses its whole data array, but operates at a higher voltage to improve reliability while sacrificing energy. We propose a hardware mechanism that allows the operating system to choose different points within that operating range based on the desired levels of performance, energy, and reliability. Alaa R. Alameldeen, Zeshan Chishti, Chris Wilkerson, Wei Wu 0024, Shih-Lien Lu |
IEEE Trans. Computers | 5 |
| 2011 | Automatic Pipelining From Transactional Datapath SpecificationsabstractThis paper presents a transactional specification framework (T-spec) for describing a datapath and the tool T-piper to synthesize automatically an in-order pipelined implementation with arbitrary user-specified pipeline-stage boundaries. T-spec abstractly views a datapath as executing one transaction at a time, computing the next system states based on the current ones. The synthesized pipeline maintains this semantics, yet allows concurrent execution of multiple overlapped transactions in different pipeline stages, where each stage performs a part of the next-state computation of each transaction. T-spec makes the state reading and writing events in a datapath explicit to enable T-piper to perform exact read-after-write (RAW) hazard analysis between the overlapped transactions. T-piper can automatically generate the pipeline control not only to ensure the correctness of the pipelined executions but also to minimize (using forwarding and speculation) the performance loss due to pipeline stalls in the presence of RAW dependencies. This paper reports design case studies applying T-spec and T-piper to reduced instruction set computing and complex instruction set computing processor pipeline development. In the latter, we report the results from a rapid design space exploration of 60 generated x86-subset pipelines, varying in pipeline depth, forwarding, and speculative execution, all starting from a single T-spec. Eriko Nurvitadhi, James C. Hoe, Timothy Kam, Shih-Lien Lu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2010 | Resilient design in scaled CMOS for energy efficiencyabstractTraditional processors are designed to guarantee error-free operation under worst-case (1) device & interconnect parameter variations resulting from less than ideal manufacturing process control; (2) static & erratic defects; (3) operating environments such as temperature excursions and voltage droops; (4) critical path activation and path delay degradations due to multiple inputs switching simultaneously in gates containing transistor stacks, or signal coupling from neighboring lines in interconnect paths; (5) speed degradation over the operating lifetime due to transistor aging under voltage, temperature & current stress; (6) early-life failures due to latent defect accelerations; and (7) soft error due to cosmic rays and alpha particle impacts. The voltage-frequency settings for all processors are set based on these infrequently encountered worst-case considerations, even though under typical conditions voltage can be pushed down further or frequency increased without causing errors for most of the processors, thus limiting both energy efficiency and performance in scaled CMOS technologies. James W. Tschanz, Keith A. Bowman, Muhammad M. Khellah, Chris Wilkerson, Bibiche M. Geuskens, Dinesh Somasekhar, Arijit Raychowdhury, Jaydeep P. Kulkarni, Carlos Tokunaga, Shih-Lien Lu, Tanay Karnik, Vivek De |
ASP-DAC | 10 |
| 2010 | Automatic multithreaded pipeline synthesis from transactional datapath specificationsabstractWe present a technique to automatically synthesize a multithreaded in-order pipeline from a high-level unpipelined datapath specification. This work extends the previously proposed transactional specification (T-spec) and synthesis technology (T-piper). The technique not only works with instruction processors but also flexible enough to accept any sequential datapath. It maintains previously proposed non-threaded pipeline features and is enhanced with multithreading features. We report a design space exploration study of 32 multithreaded x86 processor pipelines, all synthesized from a single T-spec. Eriko Nurvitadhi, James C. Hoe, Shih-Lien Lu, Timothy Kam |
DAC | 3 |
| 2010 | Automatic pipelining from transactional datapath specificationsabstractWe present a transactional datapath specification (T-spec) and the tool (T-piper) to synthesize automatically an in-order pipelined implementation from it. T-spec abstractly views a datapath as executing one transaction at a time, computing next system states based on current ones. From a T-spec, T-piper can synthesize a pipelined implementation that preserves original transaction semantics, while allowing simultaneous execution of multiple overlapped transactions across pipeline stages. T-piper not only ensures the correctness of pipelined executions, but can also employ forwarding and speculation to minimize performance loss due to data dependencies. Design case studies on RISC and CISC processor pipeline development are reported. Eriko Nurvitadhi, James C. Hoe, Timothy Kam, Shih-Lien Lu |
DATE | 4 |
| 2010 | Reducing cache power with low-cost, multi-bit error-correcting codesabstractTechnology advancements have enabled the integration of large on-die embedded DRAM (eDRAM) caches. eDRAM is significantly denser than traditional SRAMs, but must be periodically refreshed to retain data. Like SRAM, eDRAM is susceptible to device variations, which play a role in determining refresh time for eDRAM cells. Refresh power potentially represents a large fraction of overall system power, particularly during low-power states when the CPU is idle. Future designs need to reduce cache power without incurring the high cost of flushing cache data when entering low-power states. In this paper, we show the significant impact of variations on refresh time and cache power consumption for large eDRAM caches. We propose Hi-ECC, a technique that incorporates multi-bit error-correcting codes to significantly reduce refresh rate. Multi-bit error-correcting codes usually have a complex decoder design and high storage cost. Hi-ECC avoids the decoder complexity by using strong ECC codes to identify and disable sections of the cache with multi-bit failures, while providing efficient single-bit error correction for the common case. Hi-ECC includes additional optimizations that allow us to amortize the storage cost of the code over large data words, providing the benefit of multi-bit correction at same storage cost as a single-bit error-correcting (SECDED) code (2 % overhead). Our proposal achieves a 93 % reduction in refresh power vs. a baseline eDRAM cache without error correcting capability, and a 66 % reduction in refresh power vs. a system using SECDED codes. Chris Wilkerson, Alaa R. Alameldeen, Zeshan Chishti, Wei Wu 0024, Dinesh Somasekhar, Shih-Lien Lu |
ISCA | 6 |
| 2010 | Resilient microprocessor design for high performance & energy efficiencyabstractConventional microprocessors require a clock frequency (F CLK ) guardband to ensure correct functionality during infrequent dynamic operating variations in supply voltage (V CC ), temperature, and transistor aging. Consequently, these inflexible designs cannot exploit opportunities for higher performance by increasing F CLK or lower energy by reducing V CC during favorable operating conditions. This presentation describes a 45nm resilient microprocessor with error-detection and recovery circuits to detect and correct timing errors from dynamic variations to mitigate the F CLK guardband, thus enabling higher performance or lower energy as compared to a conventional design. The microprocessor core supports two distinct error-detection designs and two separate error-recovery techniques, allowing a direct comparison of the relative trade-offs. Silicon measurements demonstrate that resilient circuits enable a 41% throughput gain at equal energy or a 22% energy reduction at equal throughput, as compared to a conventional design when executing a benchmark program with a 10% V CC droop. In addition, the resilient circuits guide an adaptive clock controller that tracks recovery cycles and adapts to persistent variations by changing F CLK . The combination of error-detection and recovery circuits with dynamic adaptation allows the microprocessor to adapt to the operating environment to deliver maximum efficiency. The presentation concludes by discussing the opportunity of applying resilient techniques to enhance the dynamic operating range (i.e., high-performance and low-power modes) for microprocessors. Keith A. Bowman, James W. Tschanz, Shih-Lien Lu, Paolo A. Aseron, Muhammad M. Khellah, Arijit Raychowdhury, Bibiche M. Geuskens, Carlos Tokunaga, Chris Wilkerson, Tanay Karnik, Vivek De |
ISLPED | 3 |
| 2009 | Circuit techniques for dynamic variation toleranceabstractThree circuit techniques for dynamic variation tolerance are presented: (i) Sensors with adaptive voltage and frequency circuits, (ii) Tunable replica circuits for timing-error prediction with error recovery, and (iii) Embedded error-detection sequential circuits with error recovery. These circuits mitigate the clock frequency guardbands for dynamic variations, thus improving microprocessor performance and energy-efficiency. These circuits are described with a focus on the different trade-offs in guardband reduction and design overhead. Opportunities for CAD to further enhance microprocessor performance and energy efficiency are offered. Keith A. Bowman, James W. Tschanz, Chris Wilkerson, Shih-Lien Lu, Tanay Karnik, Vivek De, Shekhar Borkar |
DAC | 4 |
| 2009 | Resilient circuits - Enabling energy-efficient performance and reliabilityabstractVoltage and frequency margins necessary to ensure correct processor operation under dynamic voltage, temperature, and aging variations result in performance and power overheads. Resilient circuit techniques, including embedded error-detection sequentials and tunable replica circuits, allow these margins to be reduced or eliminated, resulting in reliable, energy-efficient operation. James W. Tschanz, Keith A. Bowman, Chris Wilkerson, Shih-Lien Lu, Tanay Karnik |
ICCAD | 4 |
| 2009 | Improving cache lifetime reliability at ultra-low voltagesabstractVoltage scaling is one of the most effective mechanisms to reduce microprocessor power consumption. However, the increased severity of manufacturing-induced parameter variations at lower voltages limits voltage scaling to a minimum voltage, Vccmin, below which a processor cannot operate reliably. Memory cell failures in large memory structures (e.g., caches) typically determine the Vccmin for the whole processor. Memory failures can be persistent (i.e., failures at time zero which cause yield loss) or non-persistent (e.g., soft errors or erratic bit failures). Both types of failures increase as supply voltage decreases and both need to be addressed to achieve reliable operation at low voltages. In this paper, we propose a novel adaptive technique to improve cache lifetime reliability and enable low voltage operation. This technique, multi-bit segmented ECC (MS-ECC) addresses both persistent and non-persistent failures. Like previous work on mitigating persistent failures, MS-ECC trades off cache capacity for lower voltages. However, unlike previous schemes, MS-ECC does not rely on testing to identify and isolate defective bits, and therefore enables error tolerance for nonpersistent failures like erratic bits and soft errors at low voltages. Furthermore, MS-ECC’s design can allow the operating system to adaptively change the cache size and ECC capability to adjust to system operating conditions. Compared to current designs with single-bit correction, the most aggressive implementation for MS-ECC enables a 30 % reduction in supply voltage, reducing power by 71 % and energy per instruction by 42%. Zeshan Chishti, Alaa R. Alameldeen, Chris Wilkerson, Wei Wu 0024, Shih-Lien Lu |
MICRO | 5 |
| 2008 | Novel FPGA based Haar classifier face detection algorithm accelerationabstractWe present here a novel approach to use FPGA to accelerate the Haar-classifier based face detection algorithm. With highly pipelined microarchitecture and utilizing abundant parallel arithmetic units in the FPGA, we’ve achieved real-time performance of face detection having very high detection rate and low false positives. Moreover, our approach is flexible toward the resources available on the FPGA chip. This work also provides us an understanding toward using FPGA for implementing non-systolic based vision algorithm acceleration. Our implementation is realized on a HiTech Global PCIe card that contains a Xilinx XC5VLX110T FPGA chip. Changjian Gao, Shih-Lien Lu |
FPL | 2 |
| 2008 | Trading off Cache Capacity for Reliability to Enable Low Voltage OperationabstractOne of the most effective techniques to reduce a processor's power consumption is to reduce supply voltage. However, reducing voltage in the context of manufacturing-induced parameter variations can cause many types of memory circuits to fail. As a result, voltage scaling is limited by a minimum voltage, often called Vccmin, beyond which circuits may not operate reliably. Large memory structures (e.g., caches) typically set Vccmin for the whole processor. In this paper, we propose two architectural techniques that enable microprocessor caches (L1 and L2), to operate at low voltages despite very high memory cell failure rates. The Word-disable scheme combines two consecutive cache lines, to form a single cache line where only non-failing words are used. The Bit-fix scheme uses a quarter of the ways in a cache set to store positions and fix bits for failing bits in other ways of the set. During high voltage operation, both schemes allow use of the entire cache. During low voltage operation, they sacrifice cache capacity by 50% and 25%, respectively, to reduce Vccmin below 500mV. Compared to current designs with a Vccmin of 825 mV, our schemes enable a 40% voltage reduction, which reduces power by 85% and energy per instruction (EPI) by 53%. Chris Wilkerson, Hongliang Gao, Alaa R. Alameldeen, Zeshan Chishti, Muhammad M. Khellah, Shih-Lien Lu |
ISCA | 6 |
| 2008 | A Desktop Computer with a Reconfigurable Pentium®abstractAdvancements in reconfigurable technologies, specifically FPGAs, have yielded faster, more power-efficient reconfigurable devices with enormous capacities. In our work, we provide testament to the impressive capacity of recent FPGAs by hosting a complete Pentium ® in a single FPGA chip. In addition we demonstrate how FPGAs can be used for microprocessor design space exploration while overcoming the tension between simulation speed, model accuracy, and model completeness found in traditional software simulator environments. Specifically, we perform preliminary experimentation/prototyping with an original Socket 7 based desktop processor system with typical hardware peripherals running modern operating systems such as Fedora Core 4 and Windows XP; however we have inserted a Xilinx Virtex-4 in place of the processor that should sit in the motherboard and have used the Virtex-4 to host a complete version of the Pentium ® microprocessor (which consumes less than half its resources). We can therefore apply architectural changes to the processor and evaluate their effects on the complete desktop system. We use this FPGA-based emulation system to conduct preliminary architectural experiments including growing the branch target buffer and the level 1 caches. In addition, we experimented with interfacing hardware accelerators such as DES and AES engines which resulted in a 27x speedup. Shih-Lien Lu, Peter Yiannacouras, Taeweon Suh, Rolf Kassa, Michael Konow |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2008 | Active Cache EmulatorabstractThis paper presents the active cache emulator (ACE), a novel field-programmable gate-array (FPGA)-based emulator that models an L3 cache actively and in real-time. ACE leverages interactions with its host system to model the target system. Unlike most existing FPGA-based cache emulators that collect only memory traces from their host system, ACE provides feedback to its host by injecting delays to time dilate the host system such that it experiences hit/miss latencies of the emulated cache. Such active emulation expands the context of performance evaluations by allowing measurements of system performance metrics (e.g., CPI, operations per second, frame rate) in addition to the typical cache-specific performance metrics (e.g., miss ratio) provided by existing emulators. ACE is designed to interface with a front-side bus (FSB) of a typical Pentium-based PC system. ACE utilizes the FSB snoop stall mechanism to inject delays into the system. At present, ACE is implemented using a Xilinx XC2V6000 FPGA running at 66 MHz, the same speed as its host's FSB. Verification of ACE includes using the cache calibrator and RightMark memory analyzer software to confirm proper detection of the emulated cache by the host system, and comparing ACE results with SimpleScalar software simulations. Finally, ACE is used to study L3 caches for compute-intensive, throughput-oriented, and real-time gaming benchmarks (SPEC-CPU2000, SPEC-JBB2000, Quake3). The study shows that analyzing only cache-specific metrics, as done by existing L3 cache studies with FPGA emulators, is insufficient. Active emulation mitigates this issue by providing a broader performance view, allowing researchers make better research conclusion. Eriko Nurvitadhi, Jumnit Hong, Shih-Lien Lu |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2007 | An FPGA-based Pentium in a complete desktop systemabstractSoftware simulation has been the predominant method for architects to evaluate microprocessor research proposals. There are three tenets in modeling new designs with software models: simulation speed, model accuracy and model completeness. The increasing complexity of the processor and accelerated trend to have multiple processors on a chip are putting burden on simulators to achieve all tenets mentioned, including accurately capturing OS effects. In this work we perform preliminary experimentation/prototyping with an emulation system which overcomes the tension to satisfy all three requirements. The system is an original Socket-7 based desktop processor system with typical hardware peripherals running modern operating systems such as Fedora Core 4 and Windows XP; however we have inserted a Xilinx Virtex-4 in place of the processor that should sit in the motherboard and have used the Virtex-4 to host a complete version of the Pentium® microprocessor (which consumes less than half its resources). We can therefore apply architectural changes to the processor and evaluate their effects on the complete desktop system. We use this FPGA-based emulation system to conduct preliminary architectural experiments including growing the branch target buffer and the level 1 caches. In addition, we experimented with interfacing hardware accelerators such as DES and AES engines which resulted in 27x speedups. Shih-Lien Lu, Peter Yiannacouras, Rolf Kassa, Michael Konow, Taeweon Suh |
FPGA | 1 |
| 2007 | An FPGA Approach to Quantifying Coherence Traffic Efficiency on Multiprocessor SystemsabstractRecently, there is a surge of interests in using FPGAs for computer architecture research including applications from emulating and analyzing a new platform to accelerating microarchitecural simulation speed for design space exploration. This paper proposes and demonstrates a novel usage of FPGAs for measuring the efficiency of coherent traffic of an actual computer system. Our approach employs an FPGA acting as a bus agent, interacting with a real CPU in a dual processor system to measure the intrinsic delay of coherence traffic. This technique eliminates non-deterministic factors in the measurement, such as the arbitration delay and stall in the pipelined bus. It completely isolates the impact of pure coherence traffic delay on system performance while executing workloads natively. Our experiments show that the overall execution time of the benchmark programs on a system with coherence traffic was actually increased over one without coherent traffic. It indicates that cache-to-cache transfers are less efficient in an Intel-based server system, and there exists room for further improvement such as the inclusion of the O state and cache line buffers in the memory controller. Taeweon Suh, Shih-Lien Lu, Hsien-Hsin S. Lee |
FPL | 2 |
| 2007 | Improving the reliability of on-chip data caches under process variationsabstractOn-chip caches take a large portion of the chip area. They are much more vulnerable to parameter variation than smaller units. As leakage current becomes a significant component of the total power consumption, the leakage current variations induced thermal and reliability problem to the on-chip caches become an important design concern. This paper studies the impact of process variations, particular the leakage variations, on the temperature and reliability of on-chip caches. Our statistical simulation shows that, under process variation, 85% of the caches see shortened lifetime, with average lifetime being 81.6% of the ideal cache. At runtime, unevenly distributed dynamic power and the corresponding thermal variation would further deteriorate the situation. To mitigate this problem, we propose a dynamic cache subarray permutation scheme that can alleviate the thermal stress on a high-leakage area to improve the reliability of the caches. Experiments on 17 Spec2k benchmarks show that our scheme can extend the cache lifetime by up to 20.3%, and reduce the peak temperature by 7 degrees on average and more on data-intensive applications. Wei Wu 0024, Sheldon X.-D. Tan, Jun Yang 0002, Shih-Lien Lu |
ICCD | 4 |
| 2006 | Design, implementation, and verification of active cache emulator (ACE)abstractThis paper presents the design, implementation, and verification of the Active Cache Emulator (ACE), a novel FPGA-based emulator that models an L3 cache actively and in real-time. ACE leverages interactions with its host system to model the target system (i.e. hypothetical system under study). Unlike most existing FPGA-based cache emulators that collect only memory traces from their host system, ACE provides feedback to its host by modeling the impact of the emulated cache on the system. Specifically, delays are injected to time dilate the host system which then experiences hit/miss latencies of the emulated cache. Such active emulation expands the context of performance measurements by capturing processor performance metrics (e.g. cycle per instruction) in addition to measuring the typical cache-specific performance metrics (e.g. miss ratio).ACE is designed to interface with a front-side bus (FSB) of a typical Pentium®-based PC system. To actively emulate cache latencies, ACE utilizes the snoop stall mechanism of the FSB to inject delays to the system. At present, ACE is implemented using a Xilinx XC2V6000 FPGA running at 66MHz, the same speed as its host's FSB. Verification of ACE includes using the Cache Calibrator and RightMark Memory Analyzer software to confirm proper detection of the emulated cache by the host system, and comparing ACE results with SimpleScalar software simulations. Jumnit Hong, Eriko Nurvitadhi, Shih-Lien Lu |
FPGA | 3 |
| 2006 | Research accelerator for multiple processors
David A. Patterson 0001, Arvind 0001, Krste Asanovic, Derek Chiou, James C. Hoe, Christoforos E. Kozyrakis, Shih-Lien Lu, Mark Oskin, Jan M. Rabaey, John Wawrzynek |
Hot Chips Symposium | 7 |
| 2005 | Characterization of L3 cache behavior of SPECjAppServer2002 and TPC-CabstractWith the proliferation of e-businesses, Java™ Middleware and OLTP applications are gaining importance. As the gap between CPU and memory latencies continues to increase, the performance of these applications running on multiprocessor systems will become further limited by the memory system. This study characterizes the memory behavior of such applications using the SPECjAppServer2002 and TPC-C benchmarks running on a real multiprocessor system. More specifically, the shared and private L3 caches with invalidation- and update-based coherence protocols are evaluated using the Programmable Hardware-Assisted Cache Emulator (PHA$E). We found that coherency misses increase with larger private L3 caches, constituting up to more than 15% of all misses for both benchmarks. Additionally, a saturation point was observed at which employing larger private cache yields no further improvement in miss ratio. Conversely, the shared L3 cache design was observed to be more scalable since it does not suffer from coherence misses. Our limit study shows that the existing Write-Broadcast policy, which updates line copies in other caches during a write on a shared line, has the potential to simultaneously reduce private cache miss ratio and bus traffic. For example, at 64MB, it reduces the miss ratio by 53% and 44% respectively for SPECjAppServer2002 and TPC-C, while lowering the bus traffic by 18% and 11%. In overall, the policy can eliminate the aforementioned saturation point and allows private cache miss ratio that is comparable with the miss ratio of a shared cache. Eriko Nurvitadhi, Nirut Chalainanont, Shih-Lien Lu |
ICS | 3 |
| 2003 | Implementation of HW$im - A Real-Time Configurable Cache Simulator
Shih-Lien Lu, Konrad Lai |
FPL | 1 |
| 2003 | Hardware-based Pointer Data PrefetcherabstractEffective prefetching of data from a lower memory hierarchy to higher level is a helpful way to combat the increasing memory latency that impedes the performance improvement. This paper presents a hardware-based prefetching technique to alleviate load misses caused by irregular access patterns common in linked-list structures. Different from previous works we identify a pointer load using its architecture source register. A table called target register bitmap (TRB) is maintained. By looking up this table we can identify if a load is a pointer load. We remember the base addresses of consumer load operations in a cache and prefetch the data pointed by speculative virtual addresses to a prefetch buffer, which is smaller than the data cache. Whenever a load is encountered, both the prefetch buffer and the data cache are looked up. SPEC2000 and Olden benchmarks are used to evaluate this method. This technique is able to predict accurately over 80% of the pointer load address. A system using this technique having 8KB LI data cache plus a 1KB address cache and a 1KB prefetch buffer gives an average of around 6% performance improvement over a system with 16KB LI data cache. Shih-Chang Lai, Shih-Lien Lu |
ICCD | 2 |
| 2002 | Ditto ProcessorabstractConcentration of design effort for current single-chip commercial-off-the-shelf (COTS) microprocessors has been directed towards performance. Reliability has not been the primary focus. As supply voltage scales to accommodate technology scaling and to lower power consumption, transient errors are more likely to be introduced. The basic idea behind any error tolerance scheme involves some type of redundancy. Redundancy techniques can be categorized in three general categories: (1) hardware redundancy, (2) information redundancy, and (3) time redundancy. Existing time redundant techniques for improving reliability of a superscalar processor utilize the otherwise unused hardware resources as much as possible to hide the overhead of program re-execution and verification. However, our study reveals that re-executing of long latency operations contributes to performance loss. We suggest a method to handle short and long latency instructions in slightly different ways to reduce the performance degradation. Our goal is to minimize the hardware overhead and performance degradation while maximizing the fault detection coverage. Experimental studies through microarchitecture simulation are used to compare performance lost due to the proposed scheme with non-fault tolerant design and different existing time redundant fault tolerant schemes. Fourteen integer and floating-point benchmarks are simulated with 1.8/spl sim/13.3% performance loss when compared with non-fault-tolerant superscalar processor. Shih-Chang Lai, Shih-Lien Lu, Jih-Kwon Peir |
DSN | 2 |
| 2002 | Bloom filtering cache misses for accurate data speculation and prefetchingabstractA processor must know a load instruction's latency to schedule the load's dependent instructions at the correct time. Unfortunately, modern processors do not know this latency until well after the dependent instructions should have been scheduled to avoid pipeline bubbles between themselves and the load. One solution to this problem is to predict the load's latency, by predicting whether the load will hit or miss in the data cache. Existing cache hit/miss predictors, however, can only correctly predict about 50% of cache misses.This paper introduces a new hit/miss predictor that uses a Bloom Filter to identify cache misses early in the pipeline. This early identification of cache misses allows the processor to more accurately schedule instructions that are dependent on loads and to more precisely prefetch data into the cache. Simulations using a modified SimpleScalar model show that the proposed Bloom Filter is nearly perfect, with a prediction accuracy greater than 99% for the SPECint2000 benchmarks. IPC (Instructions Per Cycle) performance improved by 19% over a processor that delayed the scheduling of instructions dependent on a load until the load latency was known, and by 6% and 7% over a processor that always predicted a load would hit the cache and with a counter-based hit/miss predictor respectively. This IPC reaches 99.7% of the IPC of a processor with perfect scheduling. Jih-Kwon Peir, Shih-Chang Lai, Shih-Lien Lu, Jared Stark, Konrad Lai |
ICS | 3 |
| 2002 | Dynamic addressing memory arrays with physical localityabstractAs pipeline width and depth grow to improve performance, memory arrays in microprocessors are growing in entries and ports. Arrays will increase in physical size, which prolongs the access time due to wiring delay. In order to boost clock frequency, these memory arrays must take multiple cycles to complete an access. This delays the scheduling of dependent instructions and affects overall performance. This paper proposes a different circuit organization to enable fast and slow accesses solely dependent on physical locality. Since the access time depends on a fixed physical location, it is pre-determined to scheduling dependent instructions. Furthermore, this paper presents a mechanism to re-configure the address decoding of the physical register file to increase the occurrence of fast accesses. Detailed circuit simulation using this proposed method determines the access cycle time. Reduction in average access cycle time for the register file and the first level data cache recovers 73% of the IPC degradation. Steven Hsu, Shih-Lien Lu, Shih-Chang Lai, Ram Krishnamurthy 0001, Konrad Lai |
MICRO | 2 |
| 2001 | Coming challenges in microarchitecture and architectureabstractIn the past several decades, the world of computers and especially that of microprocessors has witnessed phenomenal advances. Computers have exhibited ever-increasing performance and decreasing costs, making them more affordable and in turn, accelerating additional software and hardware development that fueled this process even more. The technology that enabled this exponential growth is a combination of advancements in process technology, microarchitecture, architecture, and design and development tools. While the pace of this progress has been quite impressive over the last two decades, it has become harder and harder to keep up this pace. New process technology requires more expensive megafabs and new performance levels require larger die, higher power consumption, and enormous design and validation effort. Furthermore, as CMOS technology continues to advance, microprocessor design is exposed to a new set of challenges. In the near future, microarchitecture has to consider and explicitly manage the limits of semiconductor technology, such as wire delays, power dissipation, and soft errors. In this paper we describe the role of microarchitecture in the computer world present the challenges ahead of us, and highlight areas where microarchitecture can help address these challenges. Ronny Ronen, Avi Mendelson, Konrad Lai, Shih-Lien Lu, Fred J. Pollack, John Paul Shen |
Proc. IEEE | 4 |
| 2000 | Performance improvement with circuit-level speculationabstractCurrent superscalar microprocessors' performance depends on its frequency and the number of useful instructions that can be processed per cycle (IPC). In this paper we propose a method called approximation to reduce the logic delay of a pipe-stage. The basic idea of approximation is to implement the logic function partially instead of fully. Most of the time the partial implementation gives the correct result as if the function is implemented fully but with fewer gates delay allowing a higher pipeline frequency. We apply this method on three logic blocks. Simulation results show that this method provides some performance improvement for a wide-issue superscalar if these stages are finely pipelined. Shih-Lien Lu |
MICRO | 2 |
| 1998 | Non-Stalling CounterFlow ArchitectureabstractThe counterflow pipeline concept was originated by Sproull et al.(1994) to demonstrate the concept of asynchronous circuits. This architecture relies on distributed decision making and localized clocking and data movement. We have taken these ideas and reformulated them into a substantially faster more scalable architecture that has the same distributed decision making and locality for clocking and data, but adds very aggressive speculation, no stalls, and other desirable characteristics. A high level Java simulator has been built to explore the design tradeoffs and evaluate performance. Michael F. Miller, Kenneth J. Janik, Shih-Lien Lu |
HPCA | 3 |
| 1997 | Advances of the Counterflow Pipeline MicroarchitectureabstractThe counterflow pipeline concept was originated by R.F. Sproull et al. (1994) to demonstrate the concept of asynchronous circuits. This architecture provides better throughput via clocking and data locality within the pipeline. We have taken these ideas and reformulated them into a scalable architecture that has the same locality for clocking and data, but adds aggressive speculation, fewer pipeline stalls, and a much faster startup. A high level C++ simulator has been built to explain the design tradeoffs. A VHDL model of an implementation of CFPP has been designed to validate the concept. Kenneth J. Janik, Shih-Lien Lu, Michael F. Miller |
HPCA | 2 |
| 1996 | Efficient arithmetic using self-timingabstractRecent advances in VLSI technology have facilitated high levels of integration and the implementation of faster circuits on a chip. Most of the improvements in the performance of digital systems have been brought about by such faster technologies. However, these improvements in technology have brought along with them a host of other constraints. In the faster deep submicron technologies, the wire delays constitute a significant portion of the overall delay of the system and hence some of the advantages of faster technologies are lost. The high level of integration necessitates clock distribution schemes which minimize skew across the die. These result in area penalties and adversely affect the level of integration possible at the chip level. Hence, changes in the basic architecture of computing elements of a system, which when implemented in silicon introduces reduced interconnect delays and simpler clock distribution networks, will result in more effective performance improvements. The work presented here examines the implementation of the most basic element in any datapath-an adder. The adder, a carry elimination adder (CEA), uses self-timing at both the algorithmic and implementation levels and presents a minimal hardware high speed addition mechanism. The adder exploits the nature of the input operands dynamically, which results in its average case convergence time approaching that of the ubiquitous carry lookahead adder (CLA) and the hardware complexity of a carry ripple adder (CRA). Use of self-timing results in the elimination of a global clock and hence clock-skew. Ravichandran Ramachandran, Shih-Lien Lu |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1995 | Design of a static MIMD data flow processor using micropipelinesabstractControl-flow machines are sequential in nature, executing instructions in sequence through control of program counters, whereas data-flow machines execute instructions only as input operands are made available, a process directed at the parallelism inherent within programs. At the architecture level, data-flow machines execute instructions asynchronously. In contrast, at the implementation level, the synchronous design framework of computer systems which employs globally clocked timing discipline has reached its design limits owing to problems of clock distribution. Therefore, renewed interest has been expressed in the design of computer systems based upon an asynchronous (or self-timed) approach free of the discipline imposed by the global clock. Thus, the design of a static MIMD data-flow processor using micropipelines is presented. The implemented processor, or the micro data-flow processor, differs from processors previously reported insofar as the micro data-flow processor is wholly asynchronous at both the architectural and the implementation levels.> Chih-Ming Chang, Shih-Lien Lu |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1995 | Implementation of micropipelines in enable/disable CMOS differential logicabstractThis paper examines an alternative implementation of micropipeline logic/data processing structures. To satisfy the timing requirements of the micropipeline, currently a delay element needs to be introduced in each of its stages. The alternative approach presented here eliminates this by using a differential CMOS logic family-enable/disable CMOS differential logic (ECDL) instead of the conventional static CMOS. This will ease the process of synthesizing micropipeline stages. The effectiveness of this technique in eliminating the delay requirement has been exemplified by presenting an adder implemented using ECDL.> Shih-Lien Lu |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 1988 | Device and circuit simulation interface for an integrated VLSI design environmentabstractMOSGEN, a program that provides efficient interface between the device simulator, PISCES and the circuit simulator SPICE, is described. Algorithms to generate parameters for SPICE built-in MOS transistor models have been developed and incorporated into MOSGEN. Only six PISCES simulation results are required to generate a complete set of SPICE parameters. This interface program, together with SUPREM, PISCES, and SPICE, form an integrated simulation environment for VLSI design. Such an integrated simulation environment facilitates the designers to examine just how a microscopic fabrication variable, such as implantation dose, affects final device and circuit performance as well as product yield.> Chung-Ping Wan, Bing J. Sheu, Shih-Lien Lu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |