VLDB 2026 Research / reviewers in the wild / expert
Xin Zhang 0025
dblp:76/1584-25
· DBLP profile ↗
23ranked-venue papers
3as first author
18since 2021 · last 2026
0000-0002-0579-2268ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 18 · 3 first-author · 13 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EnergAIzer: Fast and Accurate GPU Power Estimation Framework for AI WorkloadsabstractAs AI workloads drive increases in datacenter power consumption, accurate GPU power estimation is critical for proactive power management. However, existing power models face a scalability bottleneck not in the modeling techniques themselves, but in obtaining the hardware utilization inputs they require. Conventional approaches rely on either costly simulation or hardware profiling, which makes them impractical when rapid predictions are required. This work presents EnergAIzer, which addresses this scalability bottleneck by developing a lightweight solution to predict utilization inputs, reducing the estimation walltime from hours to seconds. Our key insight is that kernels in AI workloads commonly employ optimizations that create structured patterns, which analytically determine memory traffic and execution timeline. We construct a performance model using these patterns as an analytical scaffold for empirical data fitting, which also naturally exposes module-level utilization. This predicted utilization is then fed into our power model to estimate dynamic power consumption. EnergAIzer achieves $8 \%$ power errors on NVIDIA Ampere GPUs, competitive with traditional power models with elaborate cycle-level simulation or hardware profiling. We demonstrate EnergAIzer’s exploration capabilities for frequency scaling and architectural configurations, including forecasting the power of NVIDIA H100 with just $7 \%$ error. In summary, EnergAIzer provides fast and accurate power prediction for AI workloads, paving the way for power-aware design explorations. Kyungmi Lee, Zhiye Song, Xin Zhang 0025, Tamar Eilam, Anantha P. Chandrakasan |
ISPASS | 4 |
| 2026 | Extrapolation Beyond Training Data in Electric Circuits ML Tasks Using Heterogeneous-Physics-Informed GNNsabstractDue to the nonlinear nature of power converters, unpredictable operating conditions and the need to preserve physical consistency, using neural network models to predict performance beyond training data presents a significant challenge. The ability to extrapolate circuit performance lays the foundations for a robust machine learning platform for electronics design automation and circuit synthesis. In this paper, a framework for extrapolation of circuit dynamics beyond the range of the training data is proposed. To achieve this objective, a heterogeneous-physics-informed graph neural network is developed. In this method, heterogeneous graphs of electronic circuits are used to create a unique representation for each converter topology, which are then fed as inputs to a physics-informed graph neural network, thus allowing simultaneous learning of converter dynamics while enforcing the physical circuit laws. Furthermore, a detailed analysis of the impact of activation functions on mapping nonlinear circuit behavior is studied to identify and select the optimal activation function shape. The proposed approach is validated on different DC-DC converter topologies, achieving accurate interpolation and extrapolation performance. Ahmed K. Khamis, Sima Azizi Aghdam, Xin Zhang 0025, Mohammed Agamy |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2026 | Steady-State and Small-Signal Analysis of High-Ratio Hybrid Buck Converters With Enhancement to State-Space-Averaging MethodologyabstractThis paper proposes convergence enhancement to state-space averaging (SSA) methodology for steady-state and small-signal analysis of high-ratio hybrid DC-DC converters, first using analysis of Double-Step-Down (DSD) topology, including parasitics, as an example, then extending to other hybrid topologies with different numbers of capacitors and inductors. The enhanced SSA method can be used to:1)derive small-signal control-to-output transfer functions, which is essential to optimize the compensator for fast and stable closed-loop operation;2)calculate steady-state inductor currents, output voltage, input current and the voltage(s) across the flying capacitor(s),$V_{CFs}$, which is important to determine steady-state characteristics and performance;3)include circuit non-idealities such as parasitics and timing mismatches; and4)evaluate$V_{CF}$balancing property by the proposed matrix invertibility principle and added constants, and determine whether dedicated$V_{CF}$balancing circuits can be eliminated, which is considered an important benefit with reduced complexity and improved reliability. The theoretical results of DSD are then plotted in MATLAB and verified in simulations using PSIM and Cadence periodic transfer function (PXF) analysis, and measurement results using GaN devices. The simulation and measurement results match well with theoretical analysis. The enhancement is then extended beyond the DSD topology to analyze emerging hybrid topologies with more switched inductors and capacitors, future-proofing its capability to be applicable to new hybrid topologies. Muhammad Rizwan Khan, Xun Liu 0002, Xin Zhang 0025, Cheng Huang 0004 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2026 | Design of a Wide-Loading-Range Cap-Less LDO With Enhanced Power Supply Rejection and SpeedabstractThis paper introduces an output Capacitor-Less (Cap-Less) Low-Drop-Out regulator (LDO) design with NMOS pass-transistor, featuring a wide loading range with enhanced power supply rejection (PSR) and transient response. From small-signal perspective, the proposed dual-loop design achieves up to 44.6-MHz bandwidth at heavy load while maintaining stability across the full loading range from 0.1 mA to 300mA by the adaptive biasing and zero compensation techniques. From large-signal perspective, the proposed NMOS Super Source Follower (N-SSF) and small- and large-signal conflict softening filter also enhance the slew-rate for faster transient response. The proposed design is presented with bottom-up considerations and then with overall system analysis. Measurement in 180-nm CMOS shows only an 18-mV undershoot with on-chip 1 mA to 272 mA load transient steps in 100 ns with a 50-pF on-chip output capacitance. At 10 MHz, a measured PSR better than -30dB was also observed from 1-mA and 100-mA loading currents. Kejia Wang, Junyao Tang, Si Yuan Sim, Xin Zhang 0025, Cheng Huang 0004 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2026 | LaMAGIC: Advanced Circuit Formulations for Language-Model-based Topology Generation for Analog Integrated CircuitsabstractIn the realm of electronic and electrical engineering, automation of analog circuit is increasingly vital given the complexity and customized requirements of modern applications. However, existing methods only develop search-based algorithms that require many simulation iterations to design a custom circuit topology, which is usually a time-consuming process. To this end, we introduce LaMAGIC, a language model-based topology generation model that leverages supervised finetuning for automated analog circuit design. LaMAGIC can efficiently generate an optimized circuit design from the custom specification in a single pass. The generated circuit is validated by the simulator to meet the performance requirement with high precision. Our approach involves a meticulous development and analysis of various input and output formulations for circuit. These formulations can ensure canonical representations and align with the autoregressive nature of LMs for representing analog circuits as graphs. In addition, our novel transformer model supports float-input to effectively learn the mapping between numerical performance and circuits. The experimental results show that LaMAGIC achieves a success rate of up to 96% under a strict tolerance of 0.01. Also, we examine the scalability and adaptability of LaMAGIC under scarce data scenario on more complex circuits. Our findings reveal the enhanced effectiveness of our succinct float-input canonical formulation with identifier, suggesting its suitability for handling intricate circuits. Our ablation study evaluates various design choices of LM training and inference, providing insights for future domain-specific generation tasks. This research not only demonstrates the potential of language models in graph generation, but also builds a foundational framework for future explorations in automated analog circuit design. Chen-Chia Chang, Wan-Hsuan Lin, Yikang Shen, Guanglei Zhou, Yiran Chen 0001, Xin Zhang 0025 |
ACM Trans. Design Autom. Electr. Syst. | 6 |
| 2026 | Towards Generalizable and Efficient Circuit Topology Design: A Graph-Transformer-based Surrogate Model with Curriculum LearningabstractUnlike circuit parameter and sizing optimizations, the automated design of analog circuit topologies poses significant challenges for learning-based approaches. One challenge arises from the combinatorial growth of the topology space with circuit size, which limits the topology optimization efficiency. Moreover, traditional circuit evaluation methods are time-consuming, while the presence of data discontinuity in the topology space makes the accurate prediction of circuit performance exceptionally difficult for unseen topologies. To tackle these challenges, we design a novel Graph-Transformer-based Network (GTN) as the surrogate model for circuit evaluation, offering a substantial acceleration in the speed of circuit topology optimization without sacrificing performance. Our GTN model architecture is designed to embed voltage changes in circuit loops and current flows in connected devices, enabling accurate performance predictions for circuits with unseen topologies. To address the cold start problem when scaling GTN to large-scale circuits, we further introduce a curriculum learning strategy that progressively trains GTN from small-scale to large-scale circuits. This approach enables the model to first learn fundamental physical principles from simpler topologies and gradually adapt to complex configurations, effectively bridging the circuit complexity gap and improving prediction accuracy. Taking the power converter circuit design as an experimental task, our GTN model significantly outperforms an analytical approach and baseline methods directly utilizing graph neural networks. Furthermore, GTN achieves less than 5% relative error and 196× speed-up compared with high-fidelity simulation. Notably, our GTN surrogate model empowers an automatic circuit design framework to discover circuits of comparable quality to those identified through high-fidelity simulation while reducing the time required by up to 98.2%. With curriculum learning, the enhanced GTN achieves a 51% improvement for performance prediction of large-scale circuits compared to the GTN model without this strategy. These advancements establish GTN as a scalable framework for automated analog circuit design across varying circuit complexity levels. Haoshu Lu, Shaoze Fan, Ningyuan Cao, Xin Zhang 0025, Jing Li 0025 |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2025 | LaMAGIC2: Advanced Circuit Formulations for Language Model-Based Analog Topology GenerationabstractAutomation of analog topology design is crucial due to customized requirements of modern applications with heavily manual engineering efforts.
The state-of-the-art work applies a sequence-to-sequence approach and supervised finetuning on language models to generate topologies given user specifications.
However, its circuit formulation is inefficient due to $O(|V|^2)$ token length and suffers from low precision sensitivity to numeric inputs.
In this work, we introduce LaMAGIC2, a succinct float-input canonical formulation
with identifier (SFCI) for language model-based analog topology generation.
SFCI addresses these challenges by improving component-type recognition through identifier-based representations, reducing token length complexity to $O(|V|)$, and enhancing numeric precision sensitivity for better performance under tight tolerances.
Our experiments demonstrate that LaMAGIC2 achieves 34\% higher success rates under a tight tolerance 0.01 and 10X lower MSEs compared to a prior method.
LaMAGIC2 also exhibits better transferability for circuits with more vertices with up to 58.5\% improvement.
These advancements establish LaMAGIC2 as a robust framework for analog topology generation. Chen-Chia Chang, Wan-Hsuan Lin, Yikang Shen, Yiran Chen 0001, Xin Zhang 0025 |
ICML | 5 |
| 2025 | AUTOCIRCUIT-RL: Reinforcement Learning-Driven LLM for Automated Circuit Topology GenerationabstractAnalog circuit topology synthesis is integral to Electronic Design Automation (EDA), enabling the automated creation of circuit structures tailored to specific design requirements. However, the vast design search space and strict constraint adherence make efficient synthesis challenging. Leveraging the versatility of Large Language Models (LLMs), we propose AUTOCIRCUIT-RL, a novel reinforcement learning (RL)-based framework for automated analog circuit synthesis. The framework operates in two phases: instruction tuning, where an LLM learns to generate circuit topologies from structured prompts encoding design constraints, and RL refinement, which further improves the instruction-tuned model using reward models that evaluate validity, efficiency, and output voltage. The refined model is then used directly to generate topologies that satisfy the design constraints. Empirical results show that AUTOCIRCUIT-RL generates ~12% more valid circuits and improves efficiency by ~14% compared to the best baselines, while reducing duplicate generation rates by ~38%. It achieves over 60% success in synthesizing valid circuits with limited training data, demonstrating strong generalization. These findings highlight the framework's effectiveness in scaling to complex circuits while maintaining efficiency and constraint adherence, marking a significant advancement in AI-driven circuit design. Prashanth Vijayaraghavan, Luyao Shi, Ehsan Degan, Vandana V. Mukherjee, Xin Zhang 0025 |
ICML | 5 |
| 2025 | Efficient Circuit Performance Prediction Using Machine Learning: From Schematic to Layout and Silicon Measurement with Minimal Data InputabstractWe present an ML-driven framework for predicting circuit performance metrics, bridging the gap between schematic and layout simulations, multi-process corner analysis, and measured silicon data. We focus on 14nm and 5nm FinFET-based ring oscillators, collecting data across varying supply voltages, temperatures, and process corners. Using three baseline ML models—XGBoost, Random Forest, and a Neural Network—we simulate real-world design scenarios where parameter fine-tuning may not always be feasible. Key tasks include predicting layout performance from schematic data, performance prediction across process corners, and predicting measured chip performance via transfer learning. Our results show that these models can achieve less than 5% mean absolute percentage error (MAPE) for power and frequency prediction while reducing required simulations by more than 2×. In migrating from 14nm to 5nm, XGBoost and Neural Network achieve high accuracy (>0.99 R2) using just 10% of 5nm simulations. This framework offers a promising approach to accelerating circuit design across technology nodes, reducing simulation costs while maintaining accuracy in predicting performance. Dimple Vijay Kochar, Maitreyi Ashok, John Cohn, Anantha P. Chandrakasan, Xin Zhang 0025 |
ISCAS | 5 |
| 2025 | Efficient Circuit Performance Prediction Using Machine Learning: From Schematic to Layout and Silicon Measurement With Minimal Data InputabstractWe present an ML-driven framework for predicting circuit performance metrics, bridging the gap between schematic and layout simulations, multi-process corner analysis, and measured silicon data. We demonstrate this using 14nm and 5nm FinFET-based ring oscillators, by collecting data across varying supply voltages, temperatures, and process corners. Using three baseline ML models—XGBoost, Random Forest, and a Neural Network—we simulate real-world design scenarios where parameter fine-tuning may not always be feasible. Key tasks include predicting layout performance from schematic data, performance prediction across process corners, and fabricated chip performance. Our results show that these models can achieve less than 5% mean absolute percentage error (MAPE) for power and frequency prediction while reducing required simulations by more than$2\times $. When migrating from 14nm to 5nm, XGBoost and Neural Network achieve high accuracy (>0.99$R^{2}$) using just 10% of the otherwise required 5nm simulations. We also present an extensive robustness analysis to demonstrate that our results are not limited to a single data split or initialization. By varying random seeds across multiple runs, we evaluate the stability of each model with respect to algorithm initialization and the selection of training data subsets. This demonstrates that the observed accuracy is consistent and not the result of a specific, favorable configuration. This framework offers a promising approach to accelerating circuit design across technology nodes by reducing simulation costs while maintaining accuracy in predicting performance. Dimple Vijay Kochar, Maitreyi Ashok, John Cohn, Xin Zhang 0025, Anantha P. Chandrakasan |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2025 | FlexDCIM: A 400 MHz 249.1 TOPS/W 64 Kb Flexible Digital Compute-in-Memory SRAM Macro for CNN AccelerationabstractThis work proposes a 64Kb fully reconfigurable SRAM compute-in-memory (CIM) macro for convolutional neural network (CNN) acceleration using a 65nm node. It supports operation up to 400 MHz. The fully digital operation of the proposed macro effectively removes the analog CIM design issues related to process variations, noise susceptibility, and data-conversion overhead. Hence, it offers no accuracy loss, high energy efficiency, and large area saving for computation. To support the digital computation, a new area-efficient Digital Processing Unit (DPU) is proposed which is equivalent to 8.75T per bit storage. Moreover, the proposed macro features full precision reconfigurability (1b to 8b) for both input and weight, and fully flexible input activation ranging from 1 to 64 parallel inputs. It makes the proposed macro feasible for different neural network topologies. Removing sense amplifiers (SAs) for the memory mode of the proposed design suggests additional area and power savings. The proposed CIM macro achieves an energy efficiency of 249.1TOPS/W and a throughput of 819.2 GOPS. Vishal Sharma 0004, Xin Zhang 0025, Narendra Singh Dhakad, Tony Tae-Hyoung Kim |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2024 | Graph-Transformer-based Surrogate Model for Accelerated Converter Circuit Topology DesignabstractUnlike circuit parameter and sizing optimizations, the automated design of analog circuit topologies poses significant challenges for learning-based approaches. One challenge arises from the combinatorial growth of the topology space with circuit size, which limits the topology optimization efficiency. Moreover, traditional circuit evaluation methods are time-consuming, while the presence of data discontinuity in the topology space makes the accurate prediction of circuit performance exceptionally difficult for unseen topologies. To tackle these challenges, we design a novel Graph-Transformer-based Network (GTN) as the surrogate model for circuit evaluation, offering a substantial acceleration in the speed of circuit topology optimization without sacrificing performance. Our GTN model architecture is designed to embed voltage changes in circuit loops and current flows in connected devices, enabling accurate performance predictions for circuits with unseen topologies. Taking the power converter circuit design as an experimental task, our GTN model significantly outperforms an analytical approach and baseline methods directly utilizing graph neural networks. Furthermore, GTN achieves less than 5% relative error and 196× speed-up compared with high-fidelity simulation. Notably, our GTN surrogate model empowers an automatic circuit design framework to discover circuits of comparable quality to those identified through high-fidelity simulation while reducing the time required by up to 98.2%. Shaoze Fan, Haoshu Lu, Ningyuan Cao, Xin Zhang 0025, Jing Li 0025 |
DAC | 5 |
| 2024 | LaMAGIC: Language-Model-based Topology Generation for Analog Integrated CircuitsabstractIn the realm of electronic and electrical engineering, automation of analog circuit is increasingly vital given the complexity and customized requirements of modern applications. However, existing methods only develop search-based algorithms that require many simulation iterations to design a custom circuit topology, which is usually a time-consuming process. To this end, we introduce LaMAGIC, a pioneering language model-based topology generation model that leverages supervised finetuning for automated analog circuit design. LaMAGIC can efficiently generate an optimized circuit design from the custom specification in a single pass. Our approach involves a meticulous development and analysis of various input and output formulations for circuit. These formulations can ensure canonical representations of circuits and align with the autoregressive nature of LMs to effectively addressing the challenges of representing analog circuits as graphs. The experimental results show that LaMAGIC achieves a success rate of up to 96% under a strict tolerance of 0.01. We also examine the scalability and adaptability of LaMAGIC, specifically testing its performance on more complex circuits. Our findings reveal the enhanced effectiveness of our adjacency matrix-based circuit formulation with floating-point input, suggesting its suitability for handling intricate circuit designs. This research not only demonstrates the potential of language models in graph generation, but also builds a foundational framework for future explorations in automated analog circuit design. Chen-Chia Chang, Yikang Shen, Shaoze Fan, Jing Li 0025, Ningyuan Cao, Yiran Chen 0001, Xin Zhang 0025 |
ICML | 8 |
| 2024 | A 24/48V to 0.8V-1.2V All-Digital Synchronous Buck Converter with Package-Integrated GaN power FETs and 180nm Silicon Controller ICabstractThis paper presents a 24V/48V input, 0.8V- 1.2V output, two-phase, single-stage point-of-load (PoL) synchronous buck converter with enhanced-mode Gallium Nitride (GAN) N-FET based output stage and 180nm HV BCD silicon based all-digital control. The GaN devices, silicon controller chip and bootstrapping (BST) capacitors are heterogeneously integrated on an organic package substrate, thus providing a System in Package (SiP) solution, to enable high efficiency (>76% for 48:1, >86% for 24:1) while delivering 10.52W at 1V at a switching frequency of 5MHz, with a net Figure of Merit (FoM) of 11,520 MHz•V - a 13% improvement over the state-of-the art (SoA). Kaushik Bhattacharyya, Minxiang Gong, Muya Chang, Xin Zhang 0025, Arijit Raychowdhury |
ISCAS | 4 |
| 2023 | A Single-Inductor 4-Phase Hybrid Switched-Capacitor Topology for Integrated 48V-to-1V DC-DC ConvertersabstractThis paper introduces a single-inductor 4-phase hybrid switched-capacitor (4PSC) topology for integrated high-ratio direct down conversion suitable for point-of-load applications. The proposed topology consists of a 3-phase 4:1 switched-capacitor stage, reducing the switching node swing to 12-V (1/4 of the input) to significantly reduce switching loss, and an inductor to softly charge and discharge the flying capacitors with an extra phase (hence 4-phase operation) with controlled duty cycle to regulate the output voltage to 1V for direct-down conversion. The converter operates with a 4X effective switching frequency, which reduces the ripple/inductance required or switching frequency for better efficiency. With the same output voltage ripples as double step-down (DSD) and 3-level (3L1P) buck converters, the on-time is 4X that of DSD and 3L1P converters, which reduces the challenges in controller design. Lower-voltage (LV) transistors, such as 12-V devices, can be used in some of the switches to significantly improve efficiency. When compared to DSD and 3-Level converters that are state-of-the-art integrated topologies, with the same inductor, output capacitor, and output ripples in the same BCD process, this design achieves: 1) an efficiency comparable or higher than DSD (e.g., ∼3% higher at 48V-1V/5A); 2) along with using only one inductor instead of two for DSD, which can reduce the cost and increase the power density; and 3) much higher efficiency compared to a 3-level buck converter. The 4PSC topology is verified in simulations, showing peak efficiencies of ∼85% and ∼91% in a 180-nm BCD process with 48V-1V and 48V-2V conversions, respectively. Muhammad Rizwan Khan, Kang Wei 0001, Xin Zhang 0025, Cheng Huang 0004 |
ISCAS | 3 |
| 2023 | Power Converter Circuit Design Automation Using Parallel Monte Carlo Tree SearchabstractThe tidal waves of modern electronic/electrical devices have led to increasing demands for ubiquitous application-specific power converters. A conventional manual design procedure of such power converters is computation- and labor-intensive, which involves selecting and connecting component devices, tuning component-wise parameters and control schemes, and iteratively evaluating and optimizing the design. To automate and speed up this design process, we propose an automatic framework that designs custom power converters from design specifications using Monte Carlo Tree Search. Specifically, the framework embraces the upper-confidence-bound-tree (UCT), a variant of Monte Carlo Tree Search, to automate topology space exploration with circuit design specification-encoded reward signals. Moreover, our UCT-based approach can exploit small offline data via the specially designed default policy and can run in parallel to accelerate topology space exploration. Further, it utilizes a hybrid circuit evaluation strategy to substantially reduce design evaluation costs. Empirically, we demonstrated that our framework could generate energy-efficient circuit topologies for various target voltage conversion ratios. Compared to existing automatic topology optimization strategies, the proposed method is much more computationally efficient—the sequential version can generate topologies with the same quality while being up to 67% faster. The parallelization schemes can further achieve high speedups compared to the sequential version. Shaoze Fan, Ningyuan Cao, Jing Li 0025, Xin Zhang 0025 |
ACM Trans. Design Autom. Electr. Syst. | 7 |
| 2021 | From Specification to Topology: Automatic Power Converter Design via Reinforcement LearningabstractThe tidal waves of modern electronic/electrical devices have led to increasing demands for ubiquitous application-specific power converters. A conventional manual design procedure of such power converters is computation- and labor-intensive, which involves selecting and connecting component devices, tuning component-wise parameters and control schemes, and iteratively evaluating and optimizing the design. To automate and speed up this design process, we propose an automatic framework that designs custom power converters from design specifications using reinforcement learning. Specifically, the framework embraces upper-confidence-bound-tree-based (UCT-based) reinforcement learning to automate topology space exploration with circuit design specification-encoded reward signals. Moreover, our UCT-based approach can exploit small offline data via the specially designed default policy to accelerate topology space exploration. Further, it utilizes a hybrid circuit evaluation strategy to substantially reduces design evaluation costs. Empirically, we demonstrated that our framework could generate energy-efficient circuit topologies for various target voltage conversion ratios. Compared to existing automatic topology optimization strategies, the proposed method is much more computationally efficient - it can generate topologies with the same quality while being up to 67% faster. Additionally, we discussed some interesting circuits discovered by our framework. Shaoze Fan, Ningyuan Cao, Jing Li 0025, Xin Zhang 0025 |
ICCAD | 6 |
| 2021 | RaPiD: AI Accelerator for Ultra-low Precision Training and InferenceabstractThe growing prevalence and computational demands of Artificial Intelligence (AI) workloads has led to widespread use of hardware accelerators in their execution. Scaling the performance of AI accelerators across generations is pivotal to their success in commercial deployments. The intrinsic error-resilient nature of AI workloads present a unique opportunity for performance/energy improvement through precision scaling. Motivated by the recent algorithmic advances in precision scaling for inference and training, we designed RaPiD1, a 4-core AI accelerator chip supporting a spectrum of precisions, namely, 16 and 8-bit floating-point and 4 and 2-bit fixed-point. The 36mm2RaPiD chip fabricated in 7nm EUV technology delivers a peak 3.5 TFLOPS/W in HFP8 mode and 16.5 TOPS/W in INT4 mode at nominal voltage. Using a performance model calibrated to within 1% of the measurement results, we evaluated DNN inference using 4-bit fixed-point representation for a 4-core 1 RaPiD chip system and DNN training using 8-bit floating point representation for a 768 TFLOPs AI system comprising 4 32-core RaPiD chips. Our results show INT4 inference for batch size of 1 achieves 3 - 13.5 (average 7) TOPS/W and FP8 training for a mini-batch of 512 achieves a sustained 102 - 588 (average 203) TFLOPS across a wide range of applications. Swagath Venkataramani, Vijayalakshmi Srinivasan, Wei Wang 0333, Sanchari Sen, Ankur Agrawal, Monodeep Kar, Shubham Jain 0004, Alberto Mannari, Hoang Tran, Eri Ogawa, Kazuaki Ishizaki, Hiroshi Inoue, Marcel Schaal, Mauricio J. Serrano, Jungwook Choi, Xiao Sun 0013, Naigang Wang, Chia-Yu Chen, Allison Allain, James Bonanno, Nianzheng Cao, Robert Casatuta, Matthew Cohen, Bruce M. Fleischer, Michael Guillorn, Howard Haynie, Jinwook Jung, Mingu Kang, Kyu-Hyoun Kim, Siyu Koswatta, Sae Kyu Lee, Martin Lutz, Silvia M. Müller, Jinwook Oh, Ashish Ranjan 0001, Zhibin Ren, Scot Rider, Kerstin Schelm, Michael Scheuermann, Joel Silberman, Vidhi Zalani, Xin Zhang 0025, Ching Zhou, Matthew M. Ziegler, Vinay Shah, Moriyoshi Ohara, Pong-Fei Lu, Brian W. Curran, Sunil Shukla, Leland Chang, Kailash Gopalakrishnan |
ISCA | 45 |
| 2013 | A low voltage buck DC-DC converter using on-chip gate boost technique in 40nm CMOSabstractA low voltage buck DC-DC converter (0.45-V input, 0.4-V output) with on-chip gate boosted (OGB) and clock frequency scaled digital PWM controller is designed in 40-nm CMOS process. The highest efficiency to date is achieved at the output power less than 40μW. In order to compensate for the die-to-die delay variations of a delay line in the proposed digital PWM controller, a linear delay trimming by a logarithmic stress voltage (LSV) scheme with good controllability is also proposed and verified in measurement. Xin Zhang 0025, Po-Hung Chen, Yoshikatsu Ryu, Koichi Ishida, Yasuyuki Okuma, Kazunori Watanabe, Takayasu Sakurai, Makoto Takamiya |
ASP-DAC | 1 |
| 2012 | A 120-mV input, fully integrated dual-mode charge pump in 65-nm CMOS for thermoelectric energy harvesterabstractIn this paper, a fully integrated low voltage charge pump for thermoelectric energy harvesters is presented. The proposed dual-mode architecture achieves both the low startup voltage in a startup mode and high conversion efficiency in a normal operation mode without off-chip inductors and capacitors. In the measurement, the proposed circuit successfully converts 120-mV input to 770-mV output with 38.8% conversion efficiency. Po-Hung Chen, Koichi Ishida, Xin Zhang 0025, Yasuyuki Okuma, Yoshikatsu Ryu, Makoto Takamiya, Takayasu Sakurai |
ASP-DAC | 3 |
| 2012 | On-Chip Measurement System for Within-Die Delay Variation of Individual Standard Cells in 65-nm CMOSabstractNew measurement system for characterizing within-die delay variations of individual standard cells is presented. The proposed measurement system are able to characterize rising and falling delay variations separately by directly measuring the input and output waveforms of individual gate using an on-chip sampling oscilloscope in 65 nm 1.2V CMOS process. Seven types of standard cells are measured with 60 DUTs for each type. Good correlations of within-die delay distributions between measured and Monte Carlo simulated results are observed. The measured results of rising and falling delay are of great use to the modeling of standard cell library of deep-submicrometer process. By virtue of the proposed scheme, the relationship between the rising and falling delay variations and the active area of the standard cells is experimentally shown for the first time. Xin Zhang 0025, Koichi Ishida, Hiroshi Fuketa, Makoto Takamiya, Takayasu Sakurai |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2011 | An on-chip characterizing system for within-die delay variation measurement of individual standard cells in 65-nm CMOSabstractNew characterizing system for within-die delay variations of individual standard cells is presented. The proposed characterizing system is able to measure rising and falling delay variations separately by directly measuring the input and output waveforms of individual gate using an on-chip sampling oscilloscope in 65nm CMOS process. 7 types of standard cells are measured with 60 DUT's for each type. Thanks to the proposed system, a relationship between the rising and falling delay variations and the active area of the standard cells is experimentally shown for the first time. Xin Zhang 0025, Koichi Ishida, Makoto Takamiya, Takayasu Sakurai |
ASP-DAC | 1 |
| 2010 | Misleading energy and performance claims in sub/near threshold digital systemsabstractMany of us in the field of ultra-low-Vddprocessors experience difficulty in assessing the sub/near threshold circuit techniques proposed by earlier papers. This paper investigates five major pitfalls which are often not appreciated by researchers when claiming that their circuits outperform others by working at a lower Vddwith a higher energy-efficiency. These pitfalls include: i) overlook the impacts of different technologies and different Vthdefinitions, ii) only emphasize energy reduction but ignore severe throughput degradation, or expect impractical pipelining depth and parallelism degree to compensate this throughput degradation, iii) unrealistically assume that memory's Vddand energy could scale as well as standard cells, iv) use the highest temperature as the worst timing corner as in the super-threshold, but in fact negative temperature becomes much more detrimental in the sub/near threshold regime, v) pursue just-in-need Vddto compensate effects of PVT, but without considering the high energy loss on DC-DC converters. Therefore, the actual energy benefit from using a sub/near threshold Vddcan be greatly overestimated. This work provides some design guidelines and silicon evidence to ultra-low-Vddsystems. The outlined pitfalls also shed light on future directions in this field. Yu Pu, Xin Zhang 0025, Jim Huang, Atsushi Muramatsu, Masahiro Nomura, Koji Hirairi, Hidehiro Takata, Taro Sakurabayashi, Shinji Miyano, Makoto Takamiya, Takayasu Sakurai |
ICCAD | 2 |