EDBT 2026 Demo / reviewers in the wild / expert
Abbas BanaiyanMofrad
dblp:50/10300 · also Abbas Banaiyan
· DBLP profile ↗
14ranked-venue papers
7as first author
0since 2021 · last 2016
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 5 first-authorArtificial intelligence and machine learning · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 2 · 1 first-authorSecurity and privacy · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Memory systems · 46% Energy-efficient computing · 41% Hardware reliability and fault tolerance · 13% |
Topics — the 11 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems › cache › cache technology
fault-tolerant cache |
0.3 | 2 | 2015 | Power / Capacity Scaling: Energy Savings With Simple Fault-Tolerant Caches · DAC 2014 DPCS: Dynamic Power/Capacity Scaling for SRAM Caches in the Nanoscale Era · ACM Trans. Archit. Code Optim. 2015 |
Memory systems
cache |
0.2 | 1 | 2015 | DPCS: Dynamic Power/Capacity Scaling for SRAM Caches in the Nanoscale Era · ACM Trans. Archit. Code Optim. 2015 |
Energy-efficient computing › power management › memory power management
cache energy reduction |
0.2 | 1 | 2015 | DPCS: Dynamic Power/Capacity Scaling for SRAM Caches in the Nanoscale Era · ACM Trans. Archit. Code Optim. 2015 |
Energy-efficient computing › voltage scaling
dynamic voltage scaling |
0.2 | 1 | 2015 | DPCS: Dynamic Power/Capacity Scaling for SRAM Caches in the Nanoscale Era · ACM Trans. Archit. Code Optim. 2015 |
Memory systems › cache › cache technology
SRAM cache |
0.2 | 1 | 2015 | DPCS: Dynamic Power/Capacity Scaling for SRAM Caches in the Nanoscale Era · ACM Trans. Archit. Code Optim. 2015 |
Memory systems
cache design |
0.2 | 1 | 2014 | Power / Capacity Scaling: Energy Savings With Simple Fault-Tolerant Caches · DAC 2014 |
Hardware reliability and fault tolerance
memory fault tolerance |
0.2 | 1 | 2014 | Multi-Layer Memory Resiliency · DAC 2014 |
Energy-efficient computing
power gating |
0.2 | 1 | 2014 | Power / Capacity Scaling: Energy Savings With Simple Fault-Tolerant Caches · DAC 2014 |
Hardware reliability and fault tolerance
process variation |
0.1 | 1 | 2015 | DPCS: Dynamic Power/Capacity Scaling for SRAM Caches in the Nanoscale Era · ACM Trans. Archit. Code Optim. 2015 |
Energy-efficient computing
power management |
0.1 | 1 | 2014 | Power / Capacity Scaling: Energy Savings With Simple Fault-Tolerant Caches · DAC 2014 |
Energy-efficient computing
voltage scaling |
0.1 | 1 | 2014 | Power / Capacity Scaling: Energy Savings With Simple Fault-Tolerant Caches · DAC 2014 |
Methods — techniques the papers use, named apart from their topics
simulation · 0.2voltage scaling · 0.2fault disabling · 0.2architectural simulation · 0.2analytical modeling · 0.2aging mitigation · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2016 | Cross-layer virtual/physical sensing and actuation for resilient heterogeneous many-core SoCsabstractWe introduce the concepts of cross-layer virtual/physical sensing and actuation to achieve resiliency for the emerging class of heterogeneous many-core Systems-on-Chip (SoCs). Using the CyberPhysical System-on-Chip (CPSoC) concept as an exemplar sensor-rich many-core heterogeneous computing platform, we illustrate how to intrinsically couple on-chip and cross-layer physical and virtual sensing and actuation applied across different layers of the hardware/software system stack to adaptively achieve desired objectives and Quality-of-Service (QoS). We present two sample use cases that exemplify the cross-layer virtual/physical sensing and actuation approach. First, we present SmartBalance, a cross-layer sensing-driven Linux load balancer for energy efficient task execution on hetergoenous MPSOCs. Second, we present Partially Forgetful Memories, a software/hardware approach that achieves dynamic memory guard-banding for memory resilience and its application for approximate computing. Santanu Sarma, Tiago Rogério Mück, Majid Namaki-Shoushtari, Abbas BanaiyanMofrad, Nikil Dutt |
ASP-DAC | 4 |
| 2015 | Protecting caches against multi-bit errors using embedded erasure codingabstractTechnology scaling advancement coupled with operational and environmental effects make embedded memories more vulnerable to both manufacturing and transient errors including multi-bit upsets. Conventional error correcting codes incur high latency, area, and power overheads to correct multi-bit errors. In this paper, we propose Embedded Erasure Coding (EEC), a low-cost technique that can correct multi-bit errors with low overheads. This technique employs interleaved parity bits to provide a fast and low-cost multi-bit error detection. Using the erasure coding concept, the error correction is done by reconstructing the contents of the erroneous cache blocks within each cache set. Our proposed technique trades the performance for higher reliability by reserving a part of the cache (e.g. one way) to store the erasure codes. Our simulation results show that EEC provides high reliability (100% error detection and correction) with lower area overhead as compared to other state-of-the-art techniques while imposing negligible performance overhead (3%). Abbas BanaiyanMofrad, Mojtaba Ebrahimi, Fabian Oboril, Mehdi Baradaran Tahoori, Nikil Dutt |
ETS | 1 |
| 2015 | DPCS: Dynamic Power/Capacity Scaling for SRAM Caches in the Nanoscale EraabstractFault-Tolerant Voltage-Scalable (FTVS) SRAM cache architectures are a promising approach to improve energy efficiency of memories in the presence of nanoscale process variation. Complex FTVS schemes are commonly proposed to achieve very low minimum supply voltages, but these can suffer from high overheads and thus do not always offer the best power/capacity trade-offs. We observe on our 45nm test chips that the “fault inclusion property” can enable lightweight fault maps that support multiple runtime supply voltages. Based on this observation, we propose a simple and low-overhead FTVS cache architecture for power/capacity scaling. Our mechanism combines multilevel voltage scaling with optional architectural support for power gating of blocks as they become faulty at low voltages. A static (SPCS) policy sets the runtime cache VDD once such that a only a few cache blocks may be faulty in order to minimize the impact on performance. We describe a Static Power/Capacity Scaling (SPCS) policy and two alternate Dynamic Power/Capacity Scaling (DPCS) policies that opportunistically reduce the cache voltage even further for more energy savings. This architecture achieves lower static power for all effective cache capacities than a recent more complex FTVS scheme. This is due to significantly lower overheads, despite the inability of our approach to match the min-VDD of the competing work at a fixed target yield. Over a set of SPEC CPU2006 benchmarks on two system configurations, the average total cache (system) energy saved by SPCS is 62% (22%), while the two DPCS policies achieve roughly similar energy reduction, around 79% (26%). On average, the DPCS approaches incur 2.24% performance and 6% area penalties. Mark Gottscho, Abbas BanaiyanMofrad, Nikil Dutt, Alexandru Nicolau, Puneet Gupta 0001 |
ACM Trans. Archit. Code Optim. | 2 |
| 2015 | Using a Flexible Fault-Tolerant Cache to Improve Reliability for Ultra Low Voltage OperationabstractCaches are known to consume a large part of total microprocessor power. Traditionally, voltage scaling has been used to reduce both dynamic and leakage power in caches. However, aggressive voltage reduction causes process-variation--induced failures in cache SRAM arrays, which compromise cache reliability. In this article, we propose FFT-Cache, a flexible fault-tolerant cache that uses a flexible defect map to configure its architecture to achieve significant reduction in energy consumption through aggressive voltage scaling while maintaining high error reliability. FFT-Cache uses a portion of faulty cache blocks as redundancy—using block-level or line-level replication within or between sets—to tolerate other faulty caches lines and blocks. Our configuration algorithm categorizes the cache lines based on degree of conflict between their blocks to reduce the granularity of redundancy replacement. FFT-Cache thereby sacrifices a minimal number of cache lines to avoid impacting performance while tolerating the maximum amount of defects. Our experimental results on a processor executing SPEC2K benchmarks demonstrate that the operational voltage of both L1/L2 caches can be reduced down to 375 mV, which achieves up to 80% reduction in the dynamic power and up to 48% reduction in the leakage power. This comes with only a small performance loss (<%5) and 13% area overhead. Abbas BanaiyanMofrad, Houman Homayoun, Nikil Dutt |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2014 | Multi-Layer Memory ResiliencyabstractWith memories continuing to dominate the area, power, cost and performance of a design, there is a critical need to provision reliable, high-performance memory bandwidth for emerging applications. Memories are susceptible to degradation and failures from a wide range of manufacturing, operational and environmental effects, requiring a multi-layer hardware/software approach that can tolerate, adapt and even opportunistically exploit such effects. The overall memory hierarchy is also highly vulnerable to the adverse effects of variability and operational stress. After reviewing the major memory degradation and failure modes, this paper describes the challenges for dependability across the memory hierarchy, and outlines research efforts to achieve multi-layer memory resilience using a hardware/software approach. Two specific exemplars are used to illustrate multilayer memory resilience: first we describe static and dynamic policies to achieve energy savings in caches using aggressive voltage scaling combined with disabling faulty blocks; and second we show how software characteristics can be exposed to the architecture in order to mitigate the aging of large register files in GPGPUs. These approaches can further benefit from semantic retention of application intent to enhance memory dependability across multiple abstraction levels, including applications, compilers, run-time systems, and hardware platforms. Nikil Dutt, Puneet Gupta 0001, Alexandru Nicolau, Abbas BanaiyanMofrad, Mark Gottscho, Majid Namaki-Shoushtari |
DAC | 4 |
| 2014 | Power / Capacity Scaling: Energy Savings With Simple Fault-Tolerant CachesabstractComplicated approaches to fault-tolerant voltage-scalable (FTVS) SRAM cache architectures can suffer from high overheads. We propose static (SPCS) and dynamic (DPCS) variants of power/capacity scaling, a simple and low-overhead fault-tolerant cache architecture that utilizes insights gained from our 45nm SOI test chip. Our mechanism combines multi-level voltage scaling with power gating of blocks that become faulty at each voltage level. The SPCS policy sets the runtime cache VDD statically such that almost all of the cache blocks are not faulty. The DPCS policy opportunistically reduces the voltage further to save more power than SPCS while limiting the impact on performance caused by additional faulty blocks. Through an analytical evaluation, we show that our approach can achieve lower static power for all effective cache capacities than a recent complex FTVS work. This is due to significantly lower overheads, despite the failure of our approach to match the min-VDD of the competing work at fixed yield. Through architectural simulations, we find that the average energy saved by SPCS is 55%, while DPCS saves an average of 69% of energy with respect to baseline caches at 1 V. Our approach incurs no more than 4% performance and 5% area penalties in the worst case cache configuration. Mark Gottscho, Abbas BanaiyanMofrad, Nikil Dutt, Alexandru Nicolau, Puneet Gupta 0001 |
DAC | 2 |
| 2014 | NoC-based fault-tolerant cache design in chip multiprocessorsabstractAdvances in technology scaling increasingly make emerging Chip MultiProcessor (CMP) platforms more susceptible to failures that cause various reliability challenges. In such platforms, error-prone on-chip memories (caches) continue to dominate the chip area. Also, Network-on-Chip (NoC) fabrics are increasingly used to manage the scalability of these architectures. We present a novel solution for efficient implementation of fault-tolerant design of Last-Level Cache (LLC) in CMP architectures. The proposed approach leverages the interconnection network fabric to protect the LLC cache banks against permanent faults in an efficient and scalable way. During an LLC access to a faulty block, the network detects and corrects the faults, returning the fault-free data to the requesting core. Leveraging the NoC interconnection fabric, designers can implement any cache fault-tolerant scheme in an efficient, modular, and scalable manner for emerging multicore/manycore platforms. We propose four different policies for implementing a remapping-based fault-tolerant scheme leveraging the NoC fabric in different settings. The proposed policies enable design trade-offs between NoC traffic (packets sent through the network) and the intrinsic parallelism of these communication mechanisms, allowing designers to tune the system based on design constraints. We perform an extensive design space exploration on NoC benchmarks to demonstrate the usability and efficacy of our approach. In addition, we perform sensitivity analysis to observe the behavior of various policies in reaction to improvements in the NoC architecture. The overheads of leveraging the NoC fabric are minimal: on an 8-core, 16-cache-bank CMP we demonstrate reliable access to LLCs with additional overheads of less than 3% in area and less than 7% in power. Abbas BanaiyanMofrad, Gustavo Girão, Nikil Dutt |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2013 | Modeling and analysis of fault-tolerant distributed memories for networks-on-chipabstractAdvances in technology scaling increasingly make Network-on-Chips (NoCs) more susceptible to failures that cause various reliability challenges. With increasing area occupied by different on-chip memories, strategies for maintaining fault-tolerance of distributed on-chip memories become a major design challenge. We propose a system-level design methodology for scalable fault-tolerance of distributed on-chip memories in NoCs. We introduce a novel reliability clustering model for fault-tolerance analysis and shared redundancy management of on-chip memory blocks. We perform extensive design space exploration applying the proposed reliability clustering on a block-redundancy fault-tolerant scheme to evaluate the tradeoffs between reliability, performance, and overheads. Evaluations on a 64-core chip multiprocessor (CMP) with an 8x8 mesh NoC show that distinct strategies of our case study may yield up to 20% improvements in performance gains and 25% improvement in energy savings across different benchmarks, and uncover interesting design configurations. Abbas BanaiyanMofrad, Nikil Dutt, Gustavo Girão |
DATE | 1 |
| 2013 | Defuzzification block: New algorithms, and efficient hardware and software implementation issues
Hamid Reza Mahdiani, Abbas BanaiyanMofrad, Mohammad Haji Seyed Javadi, Sied Mehdi Fakhraie, Caro Lucas |
Eng. Appl. Artif. Intell. | 2 |
| 2012 | Reliable On-Chip Memory Design for CMPsabstractAggressive technology scaling in deep sub micron regime makes chips more susceptible to failures. This causes multiple realibility challenges in the design of modern chips, including manufacturing defects, wear-out, and parametric variations. With increasing area occupied by different on-chip memories in modern computing platforms such as Chip Multi-Processors (CMPs), memory reliability becomes a challenging issue. Traditional on-chip memory reliability techniques (e.g., ECC) incur significant power and performance overheads. To tackle such challenges, my research introduces several designs for fault-tolerance of both L1 and L2 cache memories in uni-core processors [1], Last-level Cache (LLC) in CMPs [3][4], and LLC in Networks-on-Chip (NoCs) [2]. Abbas BanaiyanMofrad |
SRDS | 1 |
| 2011 | FFT-cache: a flexible fault-tolerant cache architecture for ultra low voltage operationabstractCaches are known to consume a large part of total microprocessor power. Traditionally, voltage scaling has been used to reduce both dynamic and leakage power in caches. However, aggressive voltage reduction causes process-variation-induced failures in cache SRAM arrays, which compromise cache reliability. In this paper, we propose Flexible Fault-Tolerant Cache (FFT-Cache) that uses a flexible defect map to configure its architecture to achieve significant reduction in energy consumption through aggressive voltage scaling, while maintaining high error reliability. FFT-Cache uses a portion of faulty cache blocks as redundancy -- using block-level or line-level replication within or between sets to tolerate other faulty caches lines and blocks. Our configuration algorithm categorizes the cache lines based on degree of conflict of their blocks to reduce the granularity of redundancy replacement. FFT-Cache thereby sacrifices a minimal number of cache lines to avoid impacting performance while tolerating the maximum amount of defects. Our experimental results on SPEC2K benchmarks demonstrate that the operational voltage can be reduced down to 375mV, which achieves up to 80% reduction in dynamic power and up to 48% reduction in leakage power with small performance impact and area overhead. Abbas BanaiyanMofrad, Houman Homayoun, Nikil Dutt |
CASES | 1 |
| 2006 | A concurrent testing method for NoC switchesabstractThis paper proposes reuse of on-chip networks for testing switches in network on chips (NoCs). The proposed algorithm broadcasts test vectors of switches through the on-chip networks and detects faults by comparing output responses of switches with each other. This algorithm alleviates the need for: (1) external comparison of the output response of the circuit-under-test with the response of a fault free circuit stored on a tester (2) on-chip signature analysis (3) a dedicated test-bus to reach test vectors and collect their responses. Experimental results on a few test benches compare the proposed algorithm with traditional system on chip (SoC) test methods Mohammad Hosseinabady, Abbas BanaiyanMofrad, Mahdi Nazm Bojnordi, Zainalabedin Navabi |
DATE | 2 |
| 2006 | Software Implementation Issues of Existing and New Defuzzification MethodsabstractThis paper discusses software implementation issues of different defuzzification procedures in fuzzy systems. Three new defuzzification methods are introduced which are suitable for efficient software and also hardware implementations. A set of seven important existing defuzzification methods are reviewed and compared with these new methods for different software implementation approaches. The C models of all methods are prepared to perform a comprehensive analysis on the output accuracy of different methods. The results prove the superiority of our new proposed methods. In another study, three categories of low-level assembly models are developed for each method to evaluate its software execution time and instruction count when executed on each of three chosen popular processors. Namely, Texas Instruments C6xcopy DSP, Intel's Pentiumcopy IV, and IBM's PowerPC PPC405copy processors are used as the running engines for this comparison. Some accuracy-speed analysis diagrams are then introduced to guide the designers for choosing the defuzzification method which best suites their application requirements. Abbas BanaiyanMofrad, Hamid Reza Mahdiani, Sied Mehdi Fakhraie |
FUZZ-IEEE | 1 |
| 2006 | Hardware implementation and comparison of new defuzzification techniques in fuzzy processorsabstractThis paper deals with hardware implementation aspects of the defuzzification block in fuzzy controllers and processors. Three new defuzzification methods are introduced which are suitable for low cost hardware implementation. A complete set of common existing defuzzification methods are reviewed to be compared with these new methods from different hardware implementation aspects. Two different hardware models with different structures are developed for each method. The first model realizes the full combinational or fastest possible hardware implementation, and the second model demonstrates the full sequential or the smallest possible hardware implementation of each method. All models are synthesized on 0.18 micron CMOS technology cells to analyze and compare the area, delay and power consumption of different realizations of defuzzification methods. Some area-power-delay-accuracy analysis diagrams are then introduced according to the synthesis results to guide the designers through choosing the defuzzification method and also the implementation structure which best suites their application from implementation cost, speed, power consumption and output accuracy points of view Hamid Reza Mahdiani, Abbas BanaiyanMofrad, Sied Mehdi Fakhraie |
ISCAS | 2 |