EDBT 2026 Demo / reviewers in the wild / expert
Shuguang Feng
dblp:21/2673
· DBLP profile ↗
19ranked-venue papers
3as first author
0since 2021 · last 2013
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 19 · 3 first-authorSoftware engineering, systems software and programming languages · 2 · 1 first-authorSecurity and privacy · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
12 papers |
Processor architecture and microarchitecture · 26% Hardware reliability and fault tolerance · 26% Memory systems · 24% | |
| Software engineering, system software, and programming languages
3 papers |
Compilers and program optimization · 100% |
Topics — the 26 heaviest of 31, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Processor architecture and microarchitecture
chip multiprocessor |
0.5 | 5 | 2013 | Illusionist: Transforming lightweight cores into aggressive cores on demand · HPCA 2013 Erasing Core Boundaries for Robust and Configurable Performance · MICRO 2010 Necromancer: enhancing system throughput by animating dead cores · ISCA 2010 |
Memory systems › cache › cache technology
fault-tolerant cache |
0.3 | 3 | 2011 | Maximizing Spare Utilization by Virtually Reorganizing Faulty Cache Lines · IEEE Trans. Computers 2011 Archipelago: A polymorphic cache design for enabling robust near-threshold operation · HPCA 2011 ZerehCache: armoring cache architectures in high defect density technologies · MICRO 2009 |
Hardware reliability and fault tolerance
defect tolerance |
0.2 | 2 | 2011 | Maximizing Spare Utilization by Virtually Reorganizing Faulty Cache Lines · IEEE Trans. Computers 2011 Necromancer: enhancing system throughput by animating dead cores · ISCA 2010 |
Memory systems
cache design |
0.2 | 2 | 2011 | Archipelago: A polymorphic cache design for enabling robust near-threshold operation · HPCA 2011 ZerehCache: armoring cache architectures in high defect density technologies · MICRO 2009 |
Distributed systems
fault tolerance |
0.2 | 2 | 2010 | Erasing Core Boundaries for Robust and Configurable Performance · MICRO 2010 The StageNet fabric for constructing resilient multicore systems · MICRO 2008 |
Memory systems
cache |
0.1 | 1 | 2011 | Maximizing Spare Utilization by Virtually Reorganizing Faulty Cache Lines · IEEE Trans. Computers 2011 |
Processor architecture and microarchitecture › special-purpose processor
coprocessor |
0.1 | 1 | 2011 | Bundled execution of recurring traces for energy-efficient general purpose processing · MICRO 2011 |
Energy-efficient computing › energy-efficient architecture
energy-efficient processor |
0.1 | 1 | 2011 | Bundled execution of recurring traces for energy-efficient general purpose processing · MICRO 2011 |
Distributed systems › fault tolerance › resilience
graceful degradation |
0.1 | 1 | 2011 | StageNet: A Reconfigurable Fabric for Constructing Dependable CMPs · IEEE Trans. Computers 2011 |
Energy-efficient computing › voltage scaling
near-threshold voltage operation |
0.1 | 1 | 2011 | Archipelago: A polymorphic cache design for enabling robust near-threshold operation · HPCA 2011 |
Memory systems › cache design
reconfigurable cache |
0.1 | 1 | 2011 | Archipelago: A polymorphic cache design for enabling robust near-threshold operation · HPCA 2011 |
Hardware reliability and fault tolerance › error detection
transient fault detection |
0.1 | 1 | 2011 | Encore: low-cost, fine-grained transient fault recovery · MICRO 2011 |
Hardware reliability and fault tolerance › error recovery
transient fault recovery |
0.1 | 1 | 2011 | Encore: low-cost, fine-grained transient fault recovery · MICRO 2011 |
Hardware reliability and fault tolerance
soft errors |
0.1 | 1 | 2010 | Shoestring: probabilistic soft error reliability on the cheap · ASPLOS 2010 |
Hardware reliability and fault tolerance › process variation
process variation tolerance |
0.1 | 1 | 2009 | ZerehCache: armoring cache architectures in high defect density technologies · MICRO 2009 |
Processor architecture and microarchitecture › chip multiprocessor
reconfigurable multicore |
0.1 | 1 | 2008 | The StageNet fabric for constructing resilient multicore systems · MICRO 2008 |
Hardware reliability and fault tolerance
aging |
0.1 | 1 | 2007 | Self-calibrating Online Wearout Detection · MICRO 2007 |
Performance modeling and evaluation
throughput computing |
0.0 | 1 | 2013 | Illusionist: Transforming lightweight cores into aggressive cores on demand · HPCA 2013 |
Compilers and program optimization
compiler optimization |
0.0 | 1 | 2011 | Bundled execution of recurring traces for energy-efficient general purpose processing · MICRO 2011 |
Compilers and program optimization
program transformation |
0.0 | 1 | 2011 | Encore: low-cost, fine-grained transient fault recovery · MICRO 2011 |
Energy-efficient computing › voltage scaling
dynamic voltage scaling |
0.0 | 1 | 2011 | Archipelago: A polymorphic cache design for enabling robust near-threshold operation · HPCA 2011 |
Energy-efficient computing › voltage scaling
near-threshold voltage |
0.0 | 1 | 2011 | Archipelago: A polymorphic cache design for enabling robust near-threshold operation · HPCA 2011 |
Hardware reliability and fault tolerance › memory reliability
SRAM reliability |
0.0 | 1 | 2011 | Maximizing Spare Utilization by Virtually Reorganizing Faulty Cache Lines · IEEE Trans. Computers 2011 |
Electronic design automation
manufacturing yield |
0.0 | 1 | 2010 | Necromancer: enhancing system throughput by animating dead cores · ISCA 2010 |
Processor architecture and microarchitecture
pipelining |
0.0 | 1 | 2010 | Erasing Core Boundaries for Robust and Configurable Performance · MICRO 2010 |
Hardware reliability and fault tolerance
redundancy |
0.0 | 1 | 2008 | The StageNet fabric for constructing resilient multicore systems · MICRO 2008 |
Methods — techniques the papers use, named apart from their topics
trace detection · 0.2program analysis · 0.2profiling · 0.2hardware acceleration · 0.2code transformation · 0.2bundled execution · 0.2graph coloring · 0.2phase-based pruning · 0.2dynamic program distillation · 0.2minimum clique covering · 0.1instruction duplication · 0.1compile-time analysis · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2013 | Illusionist: Transforming lightweight cores into aggressive cores on demandabstractPower dissipation limits combined with increased silicon integration have led microprocessor vendors to design chip multiprocessors (CMPs) with relatively simple (lightweight) cores. While these designs provide high throughput, single-thread performance has stagnated or even worsened. Asymmetric CMPs offer some relief by providing a small number of high-performance (aggressive) cores that can accelerate specific threads. However, threads are only accelerated when they can be mapped to an aggressive core, which are restricted in number due to power and thermal budgets of the chip. Rather than using the aggressive cores to accelerate threads, this paper argues that the aggressive cores can have a multiplicative impact on single-thread performance by accelerating a large number of lightweight cores and providing an illusion of a chip full of aggressive cores. Specifically, we propose an adaptive asymmetric CMP, Illusionist, that can dynamically boost the system throughput and get a higher single-thread performance across the chip. To accelerate the performance of many lightweight cores, those few aggressive cores run all the threads that are running on the lightweight cores and generate execution hints. These hints are then used to accelerate the execution of the lightweight cores. However, the hardware resources of the aggressive core are not large enough to allow the simultaneous execution of a large number of threads. To overcome this hurdle, Illusionist performs aggressive dynamic program distillation to execute small, critical segments of each lightweight-core thread. A combination of dynamic code removal and phase-based pruning distill programs to a tiny fraction of their original contents. Experiments demonstrate that Illusionist achieves 35% higher single thread performance for all the threads running on the system, compared to a CMP with all lightweight cores, while achieving almost 2X higher system throughput compared to a CMP with all aggressive cores. Amin Ansari, Shuguang Feng, Josep Torrellas, Scott A. Mahlke |
HPCA | 2 |
| 2011 | Archipelago: A polymorphic cache design for enabling robust near-threshold operationabstractExtreme technology integration in the sub-micron regime comes with a rapid rise in heat dissipation and power density for modern processors. Dynamic voltage scaling is a widely used technique to tackle this problem when high performance is not the main concern. However, the minimum achievable supply voltage for the processor is often bounded by the large on-chip caches since SRAM cells fail at a significantly faster rate than logic cells when reducing supply voltage. This is mainly due to the higher susceptibility of the SRAM structures to process-induced parameter variations. In this work, we propose a highly flexible fault-tolerant cache design, Archipelago, that by reconfiguring its internal organization can efficiently tolerate the large number of SRAM failures that arise when operating in the near-threshold region. Archipelago partitions the cache to multiple autonomous islands with various sizes which can operate correctly without borrowing redundancy from each other. Our configuration algorithm - an adapted version of minimum clique covering - exploits the high degree of flexibility in the Archipelago architecture to reduce the granularity of redundancy replacement and minimize the amount of space lost in the cache when operating in near-threshold region. Using our approach, the operational voltage of a processor can be reduced to 375mV, which translates to 79% dynamic and 51% leakage power savings (in 90nm) for a microprocessor similar to the Alpha 21364. These power savings come with a 4.6% performance drop-off when operating in low power mode and 2% area overhead for the microprocessor. Amin Ansari, Shuguang Feng, Scott A. Mahlke |
HPCA | 2 |
| 2011 | Encore: low-cost, fine-grained transient fault recoveryabstractTo meet an insatiable consumer demand for greater performance at less power, silicon technology has scaled to unprecedented dimensions. However, the pursuit of faster processors and longer battery life has come at the cost of reliability. Given the rise of processor reliability as a first-order design constraint, there has been a growing interest in low-cost, non-intrusive techniques for transient fault detection. Many of these recent proposals have counted on the availability of hardware recovery mechanisms. Although common in aggressive out-of-order cores, hardware support for speculative rollback and recovery is less common in lower-end commodity processors. This paper presents Encore, a software-based fault recovery mechanism tailored for these lower-cost systems that lack native hardware support for speculative rollback recovery. Encore combines program analysis, profile data, and simple code transformations to create statistically idempotent code regions that can recover from faults at very little cost. Using this software-only, compiler-based approach, Encore provides the ability to recover from transient faults without specialized hardware or the costs of traditional, full-system checkpointing solutions. Experimental results show that Encore, with just 14% of runtime overhead, can safely recover, on average from 97% of transient faults when coupled with existing detection schemes. Shuguang Feng, Amin Ansari, Scott A. Mahlke, David I. August |
MICRO | 1 |
| 2011 | Bundled execution of recurring traces for energy-efficient general purpose processingabstractTechnology scaling has delivered on its promises of increasing device density on a single chip. However, the voltage scaling trend has failed to keep up, introducing tight power constraints on manufactured parts. In such a scenario, there is a need to incorporate energy-efficient processing resources that can enable more computation within the same power budget. Energy efficiency solutions in the past have typically relied on application specific hardware and accelerators. Unfortunately, these approaches do not extend to general purpose applications due to their irregular and diverse code base. Towards this end, we propose BERET, an energy-efficient co-processor that can be configured to benefit a wide range of applications. Our approach identifies recurring instruction sequences as phases of "temporal regularity" in a program's execution, and maps suitable ones to the BERET hardware, a three-stage pipeline with a bundled execution model. This judicious off-loading of program execution to a reduced-complexity hardware demonstrates significant savings on instruction fetch, decode and register file accesses energy. On average, BERET reduces energy consumption by a factor of 3-4X for the program regions selected across a range of general-purpose and media applications. The average energy savings for the entire application run was 35% over a single-issue in-order processor. Shuguang Feng, Amin Ansari, Scott A. Mahlke, David I. August |
MICRO | 2 |
| 2011 | Maximizing Spare Utilization by Virtually Reorganizing Faulty Cache LinesabstractAggressive technology scaling to 45 nm and below introduces serious reliability challenges to the design of microprocessors. Since a large fraction of chip area is devoted to on-chip caches, it is important to protect these SRAM structures against lifetime and manufacture-time failures. Designers typically overprovision caches with additional resources to overcome hard faults. However, static allocation and binding of redundant spares results in low utilization of the extra resources and ultimately limits the number of defects that can be tolerated. This work re-examines the design of process-variation-tolerant on-chip caches with a focus on providing the flexibility and dynamic reconfigurability necessary to tolerate large numbers of defects with modest hardware overhead. Our approach, ZerehCache, virtually reorganizes the cache data array using a permutation network to provide more degrees of freedom for spare allocation. A graph coloring algorithm is used to configure the network and identify the proper mapping of replacement elements. We perform an extensive design space exploration of both L1/L2 caches to identify several Pareto-optimal ZerehCaches. Given these optimal design points, we employ ZerehCache to extend the effective lifetime of the on-chip caches and prevent early lifetime failures. Finally, yield analysis studies performed on a population of 1,000 chips at the 45 nm technology node demonstrated that an L1 design with 16 percent overhead and an L2 design with eight percent area overhead achieve yields of 99 percent and 96 percent, respectively. Amin Ansari, Shuguang Feng, Scott A. Mahlke |
IEEE Trans. Computers | 3 |
| 2011 | StageNet: A Reconfigurable Fabric for Constructing Dependable CMPsabstractCMOS scaling has long been a source of dramatic performance gains. However, semiconductor feature size reduction has resulted in increasing levels of operating temperatures and current densities. Given that most wearout mechanisms are highly dependent on these parameters, significantly higher failure rates are projected for future technology generations. Consequently, fault tolerance, which has traditionally been a subject of interest for high-end server markets, is now getting emphasis in the mainstream computing systems space. The popular solution for this has been the use of redundancy at a coarse granularity, such as dual/triple modular redundancy. In this work, we challenge the practice of coarse-granularity redundancy by identifying its inability to scale to high failure rate scenarios and investigating the advantages of finer-grained configurations. To this end, this paper presents and evaluates a highly reconfigurable CMP architecture, named as StageNet (SN), that is designed with reliability as its first-class design criteria. SN relies on a reconfigurable network of replicated processor pipeline stages to maximize the useful lifetime of a chip, gracefully degrading performance toward the end of life. Our results show that the proposed SN architecture can perform 40 percent more cumulative work compared to a traditional CMP over 12 years of its lifetime. Shuguang Feng, Amin Ansari, Scott A. Mahlke |
IEEE Trans. Computers | 2 |
| 2010 | CoreGenesis: erasing core boundaries for robust and configurable performanceabstractSingle-thread performance, power efficiency and reliability are critical design challenges of future multicore systems. Although point solutions have been proposed to address these issues, a more fundamental change to the fabric of multicore systems is necessary to seamlessly combat these challenges. Towards this end, this paper proposes CoreGenesis, a dynamically adaptive multiprocessor fabric that blurs out individual core boundaries, and encourages resource sharing across cores for performance, reliability and customized processing. Further, as a manifestation of this vision, the paper provides details of a unified performance-reliability solution that can assemble variable-width processors from a network of (potentially broken) pipeline stage-level resources. Shuguang Feng, Amin Ansari, Ganesh S. Dasika, Scott A. Mahlke |
PACT | 2 |
| 2010 | Shoestring: probabilistic soft error reliability on the cheapabstractAggressive technology scaling provides designers with an ever increasing budget of cheaper and faster transistors. Unfortunately, this trend is accompanied by a decline in individual device reliability as transistors become increasingly susceptible to soft errors. We are quickly approaching a new era where resilience to soft errors is no longer a luxury that can be reserved for just processors in high-reliability, mission-critical domains. Even processors used in mainstream computing will soon require protection. However, due to tighter profit margins, reliable operation for these devices must come at little or no cost. This paper presents Shoestring, a minimally invasive software solution that provides high soft error coverage with very little overhead, enabling its deployment even in commodity processors with "shoestring" reliability budgets. Leveraging intelligent analysis at compile time, and exploiting low-cost, symptom-based error detection, Shoestring is able to focus its efforts on protecting statistically-vulnerable portions of program code. Shoestring effectively applies instruction duplication to protect only those segments of code that, when subjected to a soft error, are likely to result in user-visible faults without first exhibiting symptomatic behavior. Shoestring is able to recover from an additional 33.9% of soft errors that are undetected by a symptom-only approach, achieving an overall user-visible failure rate of 1.6%. This reliability improvement comes at a modest performance overhead of 15.8%. Shuguang Feng, Amin Ansari, Scott A. Mahlke |
ASPLOS | 1 |
| 2010 | StageWeb: Interweaving pipeline stages into a wearout and variation tolerant CMP fabricabstractManufacture-time process variation and life-time failure projections have become a major industry concern. Consequently, fault tolerance, historically of interest only for mission-critical systems, is now gaining attention in the mainstream computing space. Traditionally reliability issues have been addressed at a coarse granularity, e.g., by disabling faulty cores in chip multiprocessors. However, this is not scalable to higher failure rates. In this paper, we propose StageWeb, a fine-grained wearout and variation tolerance solution, that employs a reconfigurable web of replicated processor pipeline stages to construct dependable many-core chips. The interconnection flexibility of StageWeb simultaneously tackles wearout failures (by isolating broken stages) and process variation (by selectively disabling slower stages). Our experiments show that through its wearout tolerance, a StageWeb chip performs up to 70% more cumulative work than a comparable chip multiprocessor. Further, variation mitigation in StageWeb enables it to scale supply voltage more aggressively, resulting in up to 16% energy savings. Amin Ansari, Shuguang Feng, Scott A. Mahlke |
DSN | 3 |
| 2010 | Maestro: Orchestrating Lifetime Reliability in Chip Multiprocessors
Shuguang Feng, Amin Ansari, Scott A. Mahlke |
HiPEAC | 1 |
| 2010 | Necromancer: enhancing system throughput by animating dead coresabstractAggressive technology scaling into the nanometer regime has led to a host of reliability challenges in the last several years. Unlike on-chip caches, which can be efficiently protected using conventional schemes, the general core area is less homogeneous and structured, making tolerating defects a much more challenging problem. Due to the lack of effective solutions, disabling non-functional cores is a common practice in industry to enhance manufacturing yield, which results in a significant reduction in system throughput. Although a faulty core cannot be trusted to correctly execute programs, we observe in this work that for most defects, when starting from a valid architectural state, execution traces on a defective core actually coarsely resemble those of fault-free executions. In light of this insight, we propose a robust and heterogeneous core coupling execution scheme, Necromancer, that exploits a functionally dead core to improve system throughput by supplying hints regarding high-level program behavior. We partition the cores in a conventional CMP system into multiple groups in which each group shares a lightweight core that can be substantially accelerated using these execution hints from a potentially dead core. To prevent this undead core from wandering too far from the correct path of execution, we dynamically resynchronize architectural state with the lightweight core. For a 4-core CMP system, on average, our approach enables the coupled core to achieve 78.5% of the performance of a fully functioning core. This defect tolerance and throughput enhancement comes at modest area and power overheads of 5.3% and 8.5%, respectively. Amin Ansari, Shuguang Feng, Scott A. Mahlke |
ISCA | 2 |
| 2010 | Erasing Core Boundaries for Robust and Configurable PerformanceabstractSingle-thread performance, reliability and power efficiency are critical design challenges of future multicore systems. Although point solutions have been proposed to address these issues, a more fundamental change to the fabric of multicore systems is necessary to seamlessly combat these challenges. Towards this end, this paper proposes CoreGenesis, a dynamically adaptive multiprocessor fabric that blurs out individual core boundaries, and encourages resource sharing across cores for performance, fault tolerance and customized processing. Further, as a manifestation of this vision, the paper provides details of a unified performance-reliability solution that can assemble variable-width processors from a network of (potentially broken) pipeline stage-level resources. This design relies on interconnection flexibility, microarchitectural innovations, and compiler directed instruction steering, to merge pipeline resources for high single-thread performance. The same flexibility enables it to route around broken components, achieving sub-core level defect isolation. Together, the resulting fabric consists of a pool of pipeline stage-level resources that can be fluidly allocated for accelerating single-thread performance, throughput computing, or tolerating failures. Shuguang Feng, Amin Ansari, Scott A. Mahlke |
MICRO | 2 |
| 2009 | Adaptive online testing for efficient hard fault detectionabstractWith growing semiconductor integration, the reliability of individual transistors is expected to rapidly decline in future technology generations. In such a scenario, processors would need to be equipped with fault tolerance mechanisms to tolerate in-field silicon defects. Periodic online testing is a popular technique to detect such failures; however, it tends to impose a heavy testing penalty. In this paper, we propose an adaptive online testing framework to significantly reduce the testing overhead. The proposed approach is unique in its ability to assess the hardware health and apply suitably detailed tests. Thus, a significant chunk of the testing time can be saved for the healthy components. We further extend the framework to work with the StageNet CMP fabric, which provides the flexibility to group together pipeline stages with similar health conditions, thereby reducing the overall testing burden. For a modest 2.6% sensor area overhead, the proposed scheme was able to achieve an 80% reduction in software test instructions over the lifetime of a 16-core CMP. Amin Ansari, Shuguang Feng, Scott A. Mahlke |
ICCD | 3 |
| 2009 | Enabling ultra low voltage system operation by tolerating on-chip cache failuresabstractExtreme technology integration in the sub-micron regime comes with a rapid rise in heat dissipation and power density for modern processors. Dynamic voltage scaling is a widely used technique to tackle this problem when high performance is not needed. However, the minimum achievable supply voltage is often bounded by SRAM cells since they fail at a faster rate than logic cells. In this work, we propose a novel fault-tolerant cache architecture, that by reconfiguring its internal organization can efficiently tolerate SRAM failures that arise when operating in the ultra low voltage region. Using our approach, the operational voltage of a processor can be reduced to 420mV, which translates to 80% dynamic and 73% leakage power savings in 90nm. Amin Ansari, Shuguang Feng, Scott A. Mahlke |
ISLPED | 2 |
| 2009 | ZerehCache: armoring cache architectures in high defect density technologiesabstractAggressive technology scaling to 45nm and below introduces serious reliability challenges to the design of microprocessors. Large SRAM structures used for caches are particularly sensitive to process variation due to their high density and organization. Designers typically over-provision caches with additional resources to overcome the hard-faults. However, static allocation and binding of redundant resources results in low utilization of the extra resources and ultimately limits the number of defects that can be tolerated. This work re-examines the design of process variation tolerant on-chip caches with the focus on flexibility and dynamic reconfigurability to allow a large number defects to be tolerated with modest hardware overhead. Our approach, ZerehCache, combines redundant data array elements with a permutation network for providing a higher degree of freedom on replacement. A graph coloring algorithm is used to configure the network and find the proper mapping of replacement elements. We perform an extensive design space exploration of both L1/L2 caches to identify several Pareto optimal ZerehCaches. For the yield analysis, a population of 1000 chips was studied at the 45nm technology node; L1 designs with 16% and an L2 designs with 8% area overheads achieve yields of 99% and 96%, respectively. Amin Ansari, Shuguang Feng, Scott A. Mahlke |
MICRO | 3 |
| 2008 | StageNetSlice: a reconfigurable microarchitecture building block for resilient CMP systemsabstractAlthough CMOS feature size scaling has been the source of dramatic performance gains, it has lead to mounting reliability concerns due to increasing power densities and on-chip temperatures. Given that most wearout mechanisms that plague semiconductor devices are highly dependent on these parameters, significantly higher failure rates are projected for future technology generations. Traditional techniques for dealing with device failures have relied on coarse-grained redundancy to maintain service in the face of failed components. In this work, we challenge this practice by identifying its inability to scale to high failure rate scenarios and investigate the advantages of finer-grained configurations. We use this study to motivate the design of StageNet, an embedded CMP architecture designed from its inception with reliability as a first class design constraint. StageNet relies on a reconfigurable network of replicated processor pipeline stages to maximize the useful lifetime of the chip, gracefully degrading performance toward end of life. This paper addresses the microarchitecture of the basic building block of StageNet, named StageNetSlice, which is a processor core comprised of networked pipeline stages. A naive slice design results in approximately 4X slowdown verses a traditional processor due to longer communication delays in the pipeline. However, several small design changes that eliminate inter-stage communication paths and minimize communication bandwidth reduce this overhead to 11% on average while providing high levels of fine-grain adaptability. Shuguang Feng, Amin Ansari, Jason A. Blome, Scott A. Mahlke |
CASES | 2 |
| 2008 | The StageNet fabric for constructing resilient multicore systemsabstractScaling of CMOS feature size has long been a source of dramatic performance gains. However, the reduction in voltage levels has not been able to match this rate of scaling, leading to increasing operating temperatures and current densities. Given that most wearout mechanisms that plague semiconductor devices are highly dependent on these parameters, significantly higher failure rates are projected for future technology generations. Consequently, high reliability and fault tolerance, which have traditionally been subjects of interest for high-end server markets, are now getting emphasis in the mainstream desktop and embedded systems space. The popular solution for this has been the use of redundancy at a coarse granularity, such as dual/triple modular redundancy. In this work, we challenge the practice of coarse-granularity redundancy by identifying its inability to scale to high failure rate scenarios and investigating the advantages of finer-grained configurations. To this end, this paper presents and evaluates a highly reconfigurable multicore architecture, named StageNet (SN), that is designed with reliability as its first class design criteria. SN relies on a reconfigurable network of replicated processor pipeline stages to maximize the useful lifetime of a chip, gracefully degrading performance towards the end of life. Our results show that the proposed SN architecture can perform nearly 50% more cumulative work compared to a traditional multicore. Shuguang Feng, Amin Ansari, Jason A. Blome, Scott A. Mahlke |
MICRO | 2 |
| 2007 | Self-calibrating Online Wearout DetectionabstractTechnology scaling, characterized by decreasing feature size, thinning gate oxide, and non-ideal voltage scaling, will become a major hindrance to microprocessor reliability in future technology generations. Physical analysis of device failure mechanisms has shown that most wearout mechanisms projected to plague future technology generations are progressive, meaning that the circuit-level effects of wearout develop and intensify with age over the lifetime of the chip. This work leverages the progression of wearout over time in order to present a low-cost hardware structure that identifies increasing propagation delay, which is symptomatic of many forms of wearout, to accurately forecast the failure of microarchitectural structures. To motivate the use of this predictive technique, an HSPICE analysis of the effects of one particular failure mechanism, gate oxide breakdown, on gates from a standard cell library characterized for a 90 nm process is presented. This gate-level analysis is then used to demonstrate the aggregate change in output delay of high-level structures within a synthesized Verilog model of an embedded microprocessor core. Leveraging this analysis, a self- calibrating hardware structure for conducting statistical analysis of output delay is presented and its efficacy in predicting the failure of a variety of structures within the microprocessor core is evaluated. Jason A. Blome, Shuguang Feng, Scott A. Mahlke |
MICRO | 2 |
| 2006 | Cost-efficient soft error protection for embedded microprocessorsabstractDevice scaling trends dramatically increase the susceptibility of microprocessors to soft errors. Further, mounting demand for embedded microprocessors in a wide array of safety critical applications, ranging from automobiles to pacemakers, compounds the importance of addressing the soft error problem. Historically, soft error tolerance techniques have been targeted mainly at high-end server markets, leading to solutions such as coarse-grained modular redundancy and redundant multithreading. However, these techniques tend to be prohibitively expensive to implement in the embedded design space. To address this problem, we first present a thorough analysis of the effects of soft errors on a production-grade, fully synthesized implementation of an ARM926EJ-S embedded microprocessor. We then leverage this analysis in the design of two orthogonal low-costs of terror protection techniques that can be tuned to achieve variable levels of fault coverage as a function of area and power constraints. The first technique uses a small cache of live register values in order to provide nearly twice the fault coverage of a register file protected using traditional error correcting codes at little or no additional area cost. The second technique is a statistical method used to significantly reduce the overhead of deploying time-delayed shadow latches for low-latency fault detection. Jason A. Blome, Shuguang Feng, Scott A. Mahlke |
CASES | 3 |