Xinfei Guo

dblp:145/9137 · DBLP profile ↗
← Back
31ranked-venue papers
8as first author
24since 2021 · last 2026
0000-0002-2374-3953ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 26 · 6 first-author · 19 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CircuitS2L: Circuit Dataset Augmentation via Generative Featuring and Supervised Labeling
Linyu Zhu, Tsun-Ming Tseng, Yushan Pan, Xinfei Guo
ISCAS5
2026 Three-dimensional facility layout optimization using large language model-driven A-star and non-dominated sorting genetic Algorithm-II
Xinfei Guo, Yimiao Huang, Shaopeng Zhang
Expert Syst. Appl.1
2026 Algorithm-hardware co-design of binary neural network for efficient super resolution on FPGA
Yuanxin Su, Yushan Pan, Zhijie Xu, Xinfei Guo
Integr.7
2026 AL-HCL: Active Learning and Hierarchical Contrastive Learning for Multimodal Sentiment Analysis With Fusion Guidance
Xiaojiang He, Yushan Pan, Zhijie Xu, Zuhe Li, Xinfei Guo, Chenguang Yang 0001
IEEE Trans. Affect. Comput.5
2026 DARE: Enriching Physical Dataflow Awareness for Macro Placement Optimization
abstract
Physical dataflow, which defines the detailed connections among cells and macros, is a critical yet underexplored factor in automatic macro placement. It becomes increasingly important for enabling intelligent design automation to minimize manual intervention and reduce design iterations. Existing macro or mixed-size placers with dataflow awareness primarily focus on intrinsic relationships among macros, overlooking the crucial influence of standard cell clusters on macro placement. To address this, we propose DARE, which extracts hidden connections between macros and standard cells and incorporates a series of algorithms to enrich dataflow awareness, integrating them into placement constraints for improved macro placement. To further optimize placement results, we introduce two fine-tuning steps: (1) congestion optimization by taking macro area into consideration, and (2) flipping decisions to determine the optimal macro orientation based on the extracted dataflow information. By integrating enhanced dataflow awareness into placement constraints and applying these fine-tuning steps, the proposed approach achieves an average 7.9% improvement in half-perimeter wirelength (HPWL) across multiple widely used benchmark designs compared to a state-of-the-art dataflow-aware macro placer. Additionally, it significantly improves congestion, reducing overflow by an average of 82.5%, and achieves improvements of 36.97% in Worst Negative Slack (WNS) and 59.44% in Total Negative Slack (TNS). The approach also maintains efficient runtime throughout the entire placement, incurring less than a 1.5% runtime overhead. These results show that the proposed dataflow-driven methodology, combined with the fine-tuning steps, provides an effective foundation for macro placement within the OpenROAD flow and can be further extended to other design flows in the future to enhance placement quality.
Xiaotian Zhao, Yichen Cai 0004, Yushan Pan, Xinfei Guo
ACM Trans. Design Autom. Electr. Syst.6
2026 Guest Editorial Special Section on the International Symposium on Circuits and Systems - ISCAS 2025
Xinfei Guo, Lan-Da Van, Aatmesh Shrivastava
IEEE Trans. Very Large Scale Integr. Syst.1
2025 Revisit MBFF: Efficient Early-Stage Multi-bit Flip-Flops Clustering with Physical and Timing Awareness
abstract
Despite the maturity of Multi-bit Flip-Flops (MBFF) clustering in modern Electronic Design Automation (EDA) tools for saving power, there remains a trade-off between the flexibility to cluster flip-flops and the overall quality of results (QoR). This paper proposes a novel approach to MBFF clustering, integrating early-stage physical and timing awareness to optimize the design quality. Our pre-placement MBFF clustering algorithm addresses this trade-off by incorporating early distance estimation and predicted skews, improving timing conditions without compromising power savings. We evaluate our approach using widely-used benchmark circuits, demonstrating significant improvements in power savings and timing compared to state-of-the-art techniques. Notably, our method achieves an average improvement of 22.5% in Worst Negative Slack (WNS) and 33.91% in Total Negative Slack (TNS), while reducing power by 3.01% compared to commercial tools with MBFF clustering enabled at placement. Against the state-of-the-art pre-placement MBFF clustering algorithm, our methodology shows 25.59% and 37.97% improvements in WNS and TNS, respectively, while further reducing power by 3.08%. In addition, the proposed approach proves robust against variations in early-stage path delay estimation, maintaining superior performance even with deviations of over 20%.
Yichen Cai 0004, Linyu Zhu, Xinfei Guo
ASP-DAC3
2025 CXL-ECC: an Efficient LRC-based on-CXL-Memory-eXpander-Controller ECC to Enhance Reliability and Performance of DRAM Error Correction
abstract
Compute eXpress Link (CXL) offers an effective interface for connecting CPUs with external computing and memory devices. CXL Memory eXpander Controller (CXL-MXC) is gaining attention for its ability to boost memory capacity and bandwidth more efficiently than traditional DDR DIMMs. Despite extensive research on MXC performance and adaptation, DRAM reliability in CXL architecture remains underexplored. Traditional fault tolerance mechanisms like replica or RAID-based systems would significantly increase bandwidth overhead in the CXL fabric, adversely affecting system performance. To address this, we propose the on-CXL-Memory-Expander-Controller ECC (CXL-ECC), by using Locally Recoverable Codes (LRC) as the Inter-Channel-ECC (IC-ECC) and offloading its process to the expander, we eliminate extra memory access requests in the CXL fabric. Consequently, we conduct several experiments to demonstrate that our approach enhances DRAM reliability by more than $10^{9}$, compared to state-of-the-art ECC methods. Relative to RAID-enabled CXL switch, it reduces additional bandwidth overhead from 63.5% to 3.4% and improves system performance by 12%.
Yunfei Gu, Junhao Dai, Chentao Wu, Xinfei Guo, Jieru Zhao, Jie Li 0002, Minyi Guo
DAC6
2025 A Linear-Regression-Assisted Trimming Scheme for CMOS Voltage Reference
abstract
This work proposes a single-point linear-regression-assisted trimming method for a MOSFET-based voltage reference to reduce operation verification complexity. By obtaining the correlation between the input features at a unique temperature and the output voltages across the operating temperature range from layout-aware Monte-Carlo simulations, we can build the linear regression model to predict the profile of the voltages across the temperature range based on eight input features. Hence, we can apply the appropriate trimming code on the circuits, avoiding time-consuming temperature characterization. We design the voltage reference in 65nm CMOS and obtained 8,000 sets of simulation data to train the regression model. Validated through simulation, we reduce the temperature coefficient of the voltage reference to 64.7ppm/°C using the proposed scheme, 26% lower than that of the conventional 2-point trimming, evincing the efficiency and accuracy of the linear-regression-assisted trimming.
Chengyu Che, Xinfei Guo, Ka-Meng Lei, Rui Paulo Martins, Pui-In Mak
ISCAS3
2025 CAS-Spec: Cascade Adaptive Self-Speculative Decoding for On-the-Fly Lossless Inference Acceleration of LLMs
abstract
Speculative decoding has become a widely adopted as an effective technique for lossless inference acceleration when deploying large language models (LLMs). While on-the-fly self-speculative methods offer seamless integration and broad utility, they often fall short of the speed gains achieved by methods relying on specialized training. Cascading a hierarchy of draft models promises further acceleration and flexibility, but the high cost of training multiple models has limited its practical application. In this paper, we propose a novel Cascade Adaptive Self-Speculative Decoding (CAS-Spec) method which constructs speculative draft models by leveraging dynamically switchable inference acceleration (DSIA) strategies, including layer sparsity and activation quantization. Furthermore, traditional vertical and horizontal cascade algorithms are inefficient when applied to self-speculative decoding methods. We introduce a Dynamic Tree Cascade (DyTC) algorithm that adaptively routes the multi-level draft models and assigns the draft lengths, based on the heuristics of acceptance rates and latency prediction. Our CAS-Spec method achieves state-of-the-art acceleration compared to existing on-the-fly speculative decoding methods, with an average speedup from $1.1\times$ to $2.3\times$ over autoregressive decoding across various LLMs and datasets. DyTC improves the average speedup by $47$\% and $48$\% over cascade-based baseline and tree-based baseline algorithms, respectively. CAS-Spec can be easily integrated into most existing LLMs and holds promising potential for further acceleration as self-speculative decoding techniques continue to evolve.
Jiawei Shao, Ruge Xu, Xinfei Guo, Xuelong Li 0001
NeurIPS4
2025 Automated dimensional quality inspection of super-large steel mesh using fixed-spacing detection transformer and improved oriented fast and rotated brief
Xinfei Guo, Yimiao Huang, Shaopeng Zhang
Eng. Appl. Artif. Intell.1
2025 Back to fundamentals: Low-level visual features guided progressive token pruning
Yizhuo Liang 0002, Qingpeng Li, Xinfei Guo, Di Wu 0035, Hao Wang 0003, Yushan Pan
J. Syst. Archit.4
2025 Scale-Selectable Global Information and Discrepancy Learning Network for Multimodal Sentiment Analysis
abstract
Multimodal sentiment analysis and depression detection are pivotal for advancing human-computer interaction, yet significant challenges remain. First, the limited extraction of global contextual information within individual modalities risks the loss of modal-specific features. Second, existing methods often prioritize unaligned textual interactions, neglecting critical inter-modal discrepancies. To address these issues, we propose the Scale-Selectable Global and Discrepancy Learning Network (SSGDL), an innovative framework that integrates two core modules: the Cross-Shaped Dynamic Scale Attention Module (CSDSA) and the Primary-Secondary modal Discrepancy Learning Module (PS-MDL). The CS-DSA dynamically selects scales and employs cross-shaped attention to capture comprehensive global context and intricate internal correlations, effectively producing a fused modal representation. Meanwhile, the PS-MDL designates the fused modal as primary and utilizes cross-attention mechanisms to learn discrepancy representations between it and other modalities (textual, acoustic, and visual). By leveraging intermodal discrepancies, SSGDL achieves a more nuanced and holistic understanding of emotional content. Extensive experiments on three benchmark multimodal sentiment analysis datasets (MOSI, MOSEI, SIMS) and a depression detection dataset (AVEC2019) demonstrate that SSGDL consistently outperforms state-of-theart approaches, setting a new benchmark for multimodal affective computing.
Xiaojiang He, Yushan Pan, Xinfei Guo, Zhijie Xu, Chenguang Yang 0001
IEEE Trans. Affect. Comput.3
2025 Guest Editorial: Special Issue on the International Symposium on Integrated Circuits and Systems - ISICAS 2025
Xinfei Guo
IEEE Trans. Very Large Scale Integr. Syst.1
2024 Standard Cells Do Matter: Uncovering Hidden Connections for High-Quality Macro Placement
abstract
It becomes increasingly critical for an intelligent macro placer to be able to uncover a good dataflow for a large-scale chip to reduce churns from manual trials and errors. Existing macro placers or mixed-size placement engines that were equipped with dataflow awareness mostly focused on extracting intrinsic relations among macros only, ignoring the fact that standard cell clusters play an essential role in determining the location of macros. In this paper, we identify the necessity of macro-cell connection awareness for high-quality macro placement and propose a novel methodology to extract all “hidden” relationships efficiently among macros and cell clusters. By integrating the discovered connections as part of the placement constraints, the proposed methodology achieves an average of 2.8% and 5.5% half perimeter wire length (HPWL) improvement for considering one-hop macro-cell and two-hop macro-cell-cell dataflow connections respectively, when compared against a recently proposed dataflow-aware macro placer. A maximum of 9.7% HPWL improvement has been achieved, incurring only less than 1% runtime penalty. In addition, the congestion has been improved significantly by the proposed method, yielding an average of 62.9% and 73.4% overflow reduction for one-hop and two-hop dataflow considerations. The proposed dataflow connection extraction methodology has been demonstrated to be a significant starting point for macro placement and can be integrated into the existing design flows while delivering better design quality.
Xiaotian Zhao, Tianju Wang, Run Jiao, Xinfei Guo
DATE4
2024 One-for-All: An Unified Learning-based Framework for Efficient Cross-Corner Timing Signoff
abstract
In advanced technology nodes, the proliferation of process corners poses significant challenges in timing signoff, particularly in estimating wire-induced interconnect delay across process corners. This paper proposes a learning-based framework to perform cross-corner timing prediction efficiently and accurately. It seamlessly integrates learning-based reference corner selection and topology-aware interconnect timing prediction into the broader timing signoff steps and Engineering Change Order (ECO) processes. Unlike previous methods, it only requires information about one single known corner while accurately predicting all unknown corners. Evaluated on two mainstream industry processes, the proposed framework surpasses alternative machine learning models and existing strategies, with an impressive average accuracy enhancement of 62.6% and 95.3% respectively, and maintains a low mean absolute error (MAE) under 0.37ps and 0.01ps. Additionally, a faster version of the framework is developed to predict interconnect delay directly from a single extracted SPEF, yielding over 2× speedup. Moreover, the single-corner approach featured by the framework significantly accelerates ECO processes by over 10× compared to standard timing signoff flows. The proposed framework is also set to be open-sourced at a later date.
Linyu Zhu, Yichen Cai 0004, Xinfei Guo
ICCAD3
2024 Elastic EDA: Auto-Scaling Cloud Resources for EDA Tasks via Learning-based Approaches
abstract
Utilizing cloud EDA for chip design allows access to on-demand high-performance computing (HPC) resources, significantly reducing development time and costs by eliminating the need for costly on-site infrastructure. Despite its benefits, cloud EDA faces significant challenges, primarily the lack of an effective cost model. A key issue is the absence of a mechanism for designers to accurately gauge the characteristics of their EDA jobs in cloud environment, as the design process involves a multitude of EDA tools and steps, often leading to the over or underestimation of needed computational resources. This problem is exacerbated by the varying computational demands of different designs and constraints. To bridge this knowledge gap, we introduce Elastic EDA, a methodology that harnesses machine learning (ML) to understand the characteristics of a design and its early stages, and to predict the computational needs for subsequent phases throughout the entire EDA flow. This approach effectively aligns design behaviors with computational resources, providing cost-efficient solutions for various cloud EDA scenarios. Compared to previous ML-based predictive frameworks for cloud EDA, the proposed method achieves over 60% higher prediction accuracy and supports various elastic computing environments, maximizing the efficiency of cloud re-sources. Compared to various baseline scheduling configurations in the cloud environment, the proposed framework achieves over 16% mean runtime improvement.
Linyu Zhu, Shaogang Hao, Yushan Pan, Xinfei Guo
ICCD5
2024 CINEMA: A Configurable Binary Segmentation Based Arithmetic Module for Mixed-Precision In-Memory Acceleration
abstract
The emergence of mixed-precision quantization (MPQ) applied to edge AI models highlights the critical need for hardware support. It is a promising model compression approach with minimal accuracy loss but poses a notable hardware design challenge in the intricate balance required between computing reconfigurability and the resulting area or energy overheads. As the Compute-in-Memory (CiM) paradigm becomes prevalent for accelerating edge inference and demonstrates promising results in enhancing energy efficiency by eliminating overloaded data traffic, efficient in-memory computing circuitry becomes paramount to maximize these advantages. While the memory cell itself cannot handle complex arithmetic logic, especially in the context of MPQ-supporting mult-precision computing, the peripheral pairing module is required to be modular, portable, and scalable. In this paper, we propose a novel bit-precision configurable arithmetic module based on binary segmentation, supporting fine-grained precision ranging from 2 to 8 bits. This module achieves a peak throughput of 16 GOPS and maximum energy efficiency of 8.56 TOPS/W, occupying only 1778 μm2in 28nm technology node. The throughput and energy efficiency achieve 1.18x and 6.58x improvements compared with baseline work. Its portability enables seamless integration with various memory technologies, enhancing support for MPQ efficiently.
Runxi Wang, Ruge Xu, Xiaotian Zhao, Kai Jiang 0007, Xinfei Guo
ISCAS5
2024 Impact of Aging and Process Variability on SRAM-Based In-Memory Computing Architectures
abstract
As the SRAM-based In-memory Computing (IMC) paradigm arises as a promising candidate to break the memory wall bottleneck and deliver optimal energy efficiency, reliability has been paid less attention. Serious reliability issues that concern a conventional SRAM cell includes increased process variations as well as the dominant transistor aging mechanisms, such as Bias Temperature Instability (BTI) and Hot Carrier Injection (HCI) phenomena. These degradation impacts the speed, noise margin and input offset voltage a SRAM structure. Though the vulnerability of a SRAM cell to transistor aging has been extensively studied in the literature, there is still missing information on how aging impacts the SRAM-based IMC architectures along with the process variability. In this work, through simulation, we present a comprehensive study on impact of aging and process variability on commonly employed SRAM-based 6T-and 8T-IMC architectures. In the current study, degradation’s stochastic behavior is examined by coupling process fluctuations brought on by aging. The confluences of two degradation mechanisms help to identify the worst-case failures. This study serves as one of the first ones that provide BTI induced aging analysis on multiple IMC logic operations and offer perspectives for designing robust SRAM-based IMC based architectures.
Jani Babu Shaik, Xinfei Guo, Sonal Singhal
IEEE Trans. Circuits Syst. I Regul. Pap.2
2023 Design Space Exploration of Layer-Wise Mixed-Precision Quantization with Tightly Integrated Edge Inference Units
abstract
Layer-wise mixed-precision quantization (MPQ) has become prevailing for edge inference since it strikes a better balance between accuracy and efficiency compared to the uniform quantization scheme. Existing MPQ strategies either lacked hardware awareness or incurred huge computation costs, which gated their deployment at the edge. In this work, we propose a novel MPQ search algorithm that obtains an optimal scheme by "sampling" layer-wise sensitivity with respect to a newly proposed metric that incorporates both accuracy and proxy of hardware cost. To further efficiently deploy post-training MPQ on edge chips, we propose to tightly integrate the quantized inference units as part of the processor pipeline through micro-architecture and Instruction Set Architecture (ISA) co-design. Evaluation results show that the proposed search algorithm achieves 3% ~ 11% higher inference accuracy with similar hardware cost compared to the state-of-the-art MPQ strategies. In addition, the tightly integrated MPQ units achieve speedup of 15.13x ~ 29.65x compared to a baseline RISC-V processor.
Xiaotian Zhao, Yimin Gao, Vaibhav Verma, Ruge Xu, Mircea R. Stan, Xinfei Guo
ACM Great Lakes Symposium on VLSI6
2023 Delay-Driven Physically-Aware Logic Synthesis with Informed Search
abstract
A typical design flow is separated into front-end and back-end stages, incurring huge number of iteration loops between logic synthesis and place and route to close timing. This has been even worse in advanced technology where wire delay dominates timing. It becomes increasingly important to integrate physical awareness in the logic synthesis optimization processes to achieve better timing correlations. To tackle the physically-aware synthesis challenges, in this paper, we formulate the whole design flow as a multi-stage search problem, where informed search algorithms are utilized to perform efficient search with additional guidance. By incorporating a newly-developed learning-based routing-aware timing prediction model in the inform value function, a delay-driven physically-aware synthesis methodology called DDPAS is proposed, enabling efficient logic optimization with awareness of final routed design. The framework pairs with various search algorithms and learning-based strategies. Evaluation results show that the greedy search based DDPAS improves total negative slack (TNS) by over 58% and achieves over 7% power savings when compared to its counterpart, and the reinforcement learning (RL) based DDPAS delivers 65.6% improvement in terms of TNS and 12.1% power savings compared to a state of the art RL-based logic synthesis framework. Overall, DDPAS yields to better final QoR after routing and significantly reduces the timing closure cycles.
Linyu Zhu, Xinfei Guo
ICCD2
2023 Improving Productivity and Efficiency of SSD Manufacturing Self-Test Process by Learning-Based Proactive Defect Prediction
abstract
In the recent storage market, Flash-based Solid State Drives (SSDs) have become high-performance alternatives to Hard Disk Drives (HDDs), dramatically increasing SSD shipments. To guarantee product reliability and quality to remain competitive, SSD manufacturers pay significant efforts in technology qualification and reliability design, especially in Manufacturing Self-Test (MST) processes. However, the cost of the MST process becomes more prominent as the memory density of SSD increases. In this paper, we study the MST data in over 20,000 SSDs and propose a novel and economical approach to dynamically reduce the MST overhead by proactive infant defect prediction based on Generative Adversarial Network-Attention based Spatial-Temporal Sequence-to-Sequence network (GAN-ASTSeq). It reduces the temporal cost by 80.2% (i.e., improves the efficiency by 4×) while maintaining an outstanding detection rate of defects.
Yunfei Gu, Zixiao Chen, Chentao Wu, Xinfei Guo, Jie Li 0002, Minyi Guo, Rong Yuan, Taile Zhang, Haoran Cai
ITC5
2023 RECO-ASCON: Reconfigurable ASCON hash functions for IoT applications
Mohamed El-Hadedy 0001, Xinfei Guo, Kazutomo Yoshii, Yichen Cai 0004, Robert Herndon, Bryan Banta, Wen-Mei W. Hwu
Integr.2
2022 Agile-AES: Implementation of configurable AES primitive with agile design approach
Xinfei Guo, Mohamed El-Hadedy 0001, Sergiu Mosanu, Xiangdong Wei, Kevin Skadron, Mircea R. Stan
Integr.1
2020 Towards on-node Machine Learning for Ultra-low-power Sensors Using Asynchronous Σ Δ Streams
abstract
We propose a novel architecture to enable low-power, complex on-node data processing, for the next generation of sensors for the internet of things (IoT), smartdust, or edge intelligence. Our architecture combines near-analog-memory-computing (NAM) and asynchronous-computing-with-streams (ACS), eliminating the need for ADCs. ACS enables ultra-low power, massive computational resources required to execute on-node complex Machine Learning (ML) algorithms; while NAM addresses the memory-wall that represents a common bottleneck for ML and other complex functions. In ACS an analog value is mapped to an asynchronous stream that can take one of two logic levels ( v h , v l ). This stream-based data representation enables area/power-efficient computing units such as a multiplier implemented as an AND gate yielding savings in power of ∼90% compared to digital approaches. The generation of streams for NAM and ACS in a brute force manner, using analog-to-digital-converters (ADCs) and digital-to-streams-converters, would sky-rocket the power-latency-energy cost making the approach impractical. Our NAM-ACS architecture eliminates expensive conversions, enabling an end-to-end processing on asynchronous streams data-path. We tailor the NAM-ACS architecture for random forest (RaF), an ML algorithm, chosen for its ability to classify using a reduced number of features. Simulations show that our NAM-ACS architecture enables 75% of savings in power compared with a single ADC, obtaining a classification accuracy of 85% using an RaF-inspired algorithm.
Patricia Gonzalez-Guerrero, Tommy Tracy II, Xinfei Guo, Rahul Sreekumar, Marzieh Lenjani, Kevin Skadron, Mircea R. Stan
ACM J. Emerg. Technol. Comput. Syst.3
2019 Flexi-AES: A Highly-Parameterizable Cipher for a Wide Range of Design Constraints
abstract
Interconnected devices communicate efficiently and securely over untrusted networks via security protocols that employ various encryption algorithms, often as hardware modules. State-of-the-art hardware implementations typically focus on optimizing a single metric and are tedious to adapt to a wider set of design constraints. In this work, we develop an open-source, flexible and parameterizable hardware implementation of the Advanced Encryption Standard (AES). We present a feature-rich implementation in Chisel that is simple to employ to any architectures and to fine-tune to specific design requirements. Despite the larger design space, we use 50% fewer lines of code than existing Verilog versions, thus enabling a higher level of development productivity.
Sergiu Mosanu, Xinfei Guo, Mohamed El-Hadedy 0001, Lorena Anghel, Mircea R. Stan
FCCM2
2018 OldSpot: A Pre-RTL Model for Fine-Grained Aging and Lifetime Optimization
abstract
Modern technologies have been experiencing evergrowing power densities as they scale down, raising temperatures and increasing aging and reliability concerns. High-level simulation is necessary in order to measure aging effects on lifetime. Existing simulation tools are inadequate due to their assumptions about homogeneity and failure tolerance. These limitations reduce their accuracy in modeling heterogeneous systems or systems with shared resources, leading to lifetime overestimation. We propose a new open-source tool called "OldSpot" that relaxes these assumptions and integrates low-level models for aging mechanisms with high-level reliability modeling techniques to enable fine-grained lifetime simulation, and then show how it can be used to improve the lifetime of a multicore system by duplicating functional units rather than adding extra cores. Using this structural duplication, we show area improvement by up to 13% while eliminating performance degradation due to failing resources. Additionally, we use OldSpot to show that lifetime and temperature are not perfectly correlated and that aging must be simulated along with temperature for optimal lifetime.
Alec Roelke, Xinfei Guo, Mircea R. Stan
ICCD2
2017 Implications of accelerated self-healing as a key design knob for cross-layer resilience
Xinfei Guo, Mircea R. Stan
Integr.1
2016 Work hard, sleep well - Avoid irreversible IC wearout with proactive rejuvenation
abstract
Various wearout mechanisms have both a reversible and an irreversible (permanent) part, with some, like BTI and EM having a significant reversible part, while others, like HCI, being mostly irreversible. In this paper we make two contributions. First, we show that the boundary between the reversible and irreversible parts of wearout is not fixed, with the irreversible part becoming at least partially reversible under the right conditions of active accelerated recovery and stress/recovery scheduling. Second, we show that there are certain stress/recovery schedules that can (almost) completely eliminate irreversible wearout, thus allowing significant reductions in necessary design margins. The experiments were done on commercial FPGAs fabricated in a 40nm technology. To fully repair and avoid the irreversible wearout, we propose a biology-inspired sleep-when-getting-tired strategy. The strategy can achieve >60× design margin reduction and ~9% average performance improvement within a 10-year lifetime constraint compared to the no-recovery case. Potential system level implementations (a negative “turbo-boost” like strategy) in multicore and NoC systems are also presented.
Xinfei Guo, Mircea R. Stan
ASP-DAC1
2014 Modeling and Experimental Demonstration of Accelerated Self-Healing Techniques
abstract
In this paper we postulate that future electronics systems will use sleep time as an active recovery period essential for their overall performance. Our hypothesis is that by explicitly controlling the ratio of sleep vs. active and sleep conditions (e.g. higher temperatures, negative voltages), we can deeply rejuvenate electronic systems periodically to improve their metrics. We perform a series of stress and recovery experiments using commercial FPGAs to demonstrate several cases where we bring stressed chips back to within 90% of their original margin by actively rejuvenating for only 1/4 of the stress time. We validate our experiments against extracted models and present potential applications to multi-core systems.
Xinfei Guo, Wayne P. Burleson, Mircea R. Stan
DAC1
2014 A multi-output on-chip switched-capacitor DC-DC converter for near- and sub-threshold power modes
abstract
This paper presents a novel multi-output on-chip switched-capacitor (SC) DC-DC converter simultaneously providing two output voltages (2VDD/3 and VDD/3, in addition to the available full VDD) that enables the use of three different power modes for optimizing power/performance trade-offs: super-threshold (full VDD), near-threshold (2VDD/3) and sub-threshold (VDD/3). Unlike previously proposed SC converters, the multiple conversion ratios of the novel converter are achieved without changing the topology of the circuit, thus fewer components being needed to support multiple voltages. Above 84% and 70% efficiencies are obtained in simulation over a range of load currents from 0.4 mA to 5 mA, and from 0.4 mA to 1 mA, with conversion ratios of 2/3 and 1/3, respectively. The paper also presents a novel optimization method to save area by using unequally sized flying capacitors. The target application is for systems with discrete power modes; also for systems that use dithering to emulate a continuous range of voltages such as the Panoptic Dynamic Voltage Scaling (PDVS).
Yingbo Zhao, Yintang Yang, Kaushik Mazumdar, Xinfei Guo, Mircea R. Stan
ISCAS4