Ian A. Young

dblp:39/6837 · also Ian Young 0001 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
5since 2021 · last 2026
0000-0002-4017-5265ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Ultra Memory-Efficient On-FPGA Training of Transformers via Tensor-Compressed Optimization
abstract
Transformer models have achieved state-of-the-art performance across a wide range of machine learning tasks. There is growing interest in training transformers on resource-constrained edge devices due to considerations such as privacy, domain adaptation, and on-device scientific machine learning. However, the significant computational and memory demands required for transformer training often exceed the capabilities of an edge device. Leveraging low-rank tensor compression, this paper presents the first on-FPGA accelerator for transformer training. On the algorithm side, we present a bi-directional contraction flow for tensorized transformer training, significantly reducing the computational FLOPS and intra-layer memory costs compared to existing tensor operations. On the hardware side, we store all highly compressed model parameters and gradient information on chip, creating an on-chip-memory-only framework for each stage in training. This reduces off-chip communication and minimizes latency and energy costs. Additionally, we implement custom computing kernels for each training stage and employ intra-layer parallelism and pipe-lining to further enhance run-time and memory efficiency. Through experiments on transformer models within 36.7 to 93.5 MB using FP-32 data formats on the ATIS dataset, our tensorized FPGA accelerator could conduct single-batch end-to-end training on the AMD Alevo U50 FPGA, with a memory budget of less than 6-MB BRAM and 22.5-MB URAM. Compared to uncompressed training on the NVIDIA RTX 3090 GPU, our on-FPGA training achieves a memory reduction of 30× to 51×. Our FPGA accelerator also achieves up to 4.0× less energy cost per epoch compared with tensor transformer training on an NVIDIA RTX 3090 GPU. As an initial result, this work highlights the significant potential of large-scale tensor training on edge devices.
Jinming Lu, Hai Li 0008, Cong Hao, Ian A. Young, Zheng Zhang 0005
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2025 Enhanced Operator Learning for Scalable and Ultra-fast Thermal Simulation in 3D-IC Design
abstract
Thermal simulation plays a critical role in the design of 3D integrated circuits (3D-ICs), where accurate and efficient temperature predictions are essential to ensure component reliability and performance. Recently, deep learning methods have shown great potential in accelerating these simulations. DeepOHeat [1] is one such approach, designed to learn solution operators that map single or multiple configurations of heat equations---such as surface power, volume power, or heat transfer coefficients---directly to the 3D temperature distribution. By utilizing a physics-informed DeepONet framework[2], DeepOHeat effectively captures complex relationships between design parameters and temperature fields, even when no training data is available and only PDE constraints are imposed.
Xinling Yu, Ziyue Liu 0003, Hai Li 0008, Ian A. Young, Zheng Zhang 0005
ASP-DAC4
2025 Digital Compute-in-Memory Ising Annealer with Ferroelectric Capacitor-Based nvSRAM for Combinatorial Optimization Problems
abstract
Combinatorial optimization problems (COPs) have a wide range of applications. The Ising model-based annealer is gaining attention for its efficiency and speed in finding approximate solutions. However, building an Ising machine that is area- and energy-efficient, scalable, and with low compute latency in CMOS is challenging. In this paper, we present a digital compute-in-memory (DCIM) Ising annealer that uses ferroelectric capacitor (FeCap)-based nvSRAM to solve COPs like the Traveling Salesman Problem (TSP). By using weak recall operations, our design eliminates the need to reload weights, significantly reducing energy consumption and speeding up processing compared to other approaches. Simulations using a 16nm PDK demonstrate that our nvSRAM-based DCIM array maintains accuracy while reducing latency by up to 55.0% and energy by 49.6% compared to prior work implemented with conventional SRAM DCIM array. Algorithm validation further shows that the random noise introduced by weak recall can be effectively utilized in the annealing process.
Yuyao Kong, Jianwei Jia, Anni Lu, Faaiq G. Waqar, Yuan-Chun Luo, Hai Li 0001, Ian A. Young, Shimeng Yu
ISCAS7
2025 Poor Man's Training on MCUs: A Memory-Efficient Quantized Back-Propagation-Free Approach
abstract
Back propagation (BP) is the default solution for gradient computation in neural network training. However, implementing BP-based training on various edge devices such as FPGA, microcontrollers (MCUs), and analog computing platforms faces multiple major challenges, such as the lack of hardware resources, long time-to-market, and dramatic errors in a low-precision setting. This article presents a simple BP-free training scheme on an MCU, which makes edge training hardware design as easy as inference hardware design. We adopt a quantized zeroth-order method to estimate the gradients of quantized model parameters, which can overcome the error of a straight-through estimator in a low-precision BP scheme. We further employ a few dimension reduction methods (e.g., node perturbation, sparse training) to improve the convergence of zeroth-order training. Experiment results show that our BP-free training achieves comparable performance as BP-based training on adapting a pre-trained image classifier to various corrupted data on resource-constrained edge devices (e.g., an MCU with 1024-KB SRAM for dense full-model training, or an MCU with 256-KB SRAM for sparse training). This method is most suitable for application scenarios where memory cost and time-to-market are the major concerns, but longer latency can be tolerated.
Yequan Zhao, Hai Li 0008, Ian A. Young, Zheng Zhang 0005
ACM Trans. Design Autom. Electr. Syst.3
2024 Digital CIM with Noisy SRAM Bit: A Compact Clustered Annealer for Large-Scale Combinatorial Optimization
abstract
Combinatorial optimization problems (COP) are NP-hard and intractable to solve using conventional computing. The Ising model-based annealer has gained increasing attention recently due to its efficiency and speed in finding approximate solutions. However, Ising solvers for travelling salesman problems (TSP) usually suffer from a scalability issue due to quadratically increasing number of spins. In this paper, we propose a digital computing-in-memory (CIM) based clustered annealer to solve tens of thousands of city-scale TSP with only a few mega-byte (MB) of static random access memory (SRAM), using hierarchical clustering to solve input sparsity and digital CIM flexibility to solve weight sparsity. The intrinsic process variations between SRAM devices are utilized to generate the noisy bit errors during pseudo-read under reduced supply voltage, realizing the annealing process. The design space of cluster size and programmability is explored to understand the trade-offs of solution quality and hardware cost, for TSP scale ranging from 3080 to 85900 cities. The proposed design speeds up the convergence by >109× with <25% solution quality overhead compared with the CPU baseline. The comparison with state-of-the-art scalable annealers shows a >1013× improvement on functionally normalized area and power.
Anni Lu, Yuan-Chun Luo, Hai Li 0008, Ian A. Young, Shimeng Yu
DAC5
2019 An Energy-Efficient Classifier via Boosted Spin Channel Networks
abstract
With diminishing energy and delay benefits via CMOS scaling, there is much interest in exploring the use of alternative state variables such as electronic spin. Multiple research efforts are underway exploring both Boolean and non-Boolean design space using spin devices in order to make their energy and delay benefits competitive to CMOS. In this paper, we propose spin channel networks (SCNs) - spin-based circuits that exploit exponential decay of spin current to efficiently realize multi-bit dot product computation. We show that proposed SCNs can be employed with adaptive boosting (AdaBoost) learning algorithm to efficiently realize a binary classifier for breast cancer detection. The proposed SCN implementation achieves 112× and 14× lower energy per decision compared to the conventional all spin logic (ASL) and 20nm CMOS designs, respectively, for identical decision throughput.
Ameya Patil 0001, Sasikanth Manipatruni, Dmitri E. Nikonov, Ian A. Young, Naresh R. Shanbhag
ISCAS4
2018 Density Tradeoffs of Non-Volatile Memory as a Replacement for SRAM Based Last Level Cache
abstract
Increasing the capacity of the Last Level Cache (LLC) can help scale the memory wall. Due to prohibitive area and leakage power, however, growing conventional SRAM LLC already incurs diminishing returns. Emerging Non-Volatile Memory (NVM) technologies like Spin Torque Transfer RAM (STTRAM) promise high density and low leakage, thereby offering an attractive alternative for building large capacity LLCs. However these technologies have significantly longer write latency compared to SRAM, which interferes with reads and severely limits their performance potential. Despite the recent work showing the write latency reduction at NVM technology level, practical considerations like high yield and low bit error rates will result a significant loss of NVM density when these techniques are implemented. Therefore, improving the write latency while compromising on the density results in sub-optimal usage of the NVM technology. In this paper we present a novel STTRAM LLC design that mitigates the long write latency, thereby delivering SRAM like performance while preserving the benefits of high density. Based on a light-weight learning mechanism, our solution relieves LLC congestion through two schemes. Firstly, we propose write congestion aware bypass that eliminates a large fraction of writes. Despite dropping LLC hit rates which could severely degrade performance in a conventional LLC, our policy smartly modulates the bypass, overcomes the hit rate loss and delivers significant performance gain. Furthermore, our solution establishes a virtual hybrid cache that absorbs and eliminates the redundant writes, which otherwise might be repeatedly and slowly written to the NVM LLC. Detailed simulation of traditional SPEC CPU 2006 suite as well as important industry workloads running on a 4-core system shows that our proposal delivers on an average 26% performance improvement over a baseline LLC design using 8MB STTRAM, while reducing the memory system energy by 10%. Our design outperforms a similar area SRAM LLC by nearly 18%, thereby making NVM technology an attractive alternative for future high performance computing.
Kunal Korgaonkar, Ishwar Bhati, Huichu Liu, Jayesh Gaur, Sasikanth Manipatruni, Sreenivas Subramoney, Tanay Karnik, Steven Swanson, Ian A. Young, Hong Wang 0003
ISCA9
2017 A Systems Approach to Computing in Beyond CMOS Fabrics: Invited
abstract
No abstract available.
Ameya Patil 0001, Naresh R. Shanbhag, Lav R. Varshney, Eric Pop, H.-S. Philip Wong, Subhasish Mitra, Jan M. Rabaey, Jeffrey A. Weldon, Lawrence T. Pileggi, Sasikanth Manipatruni, Dmitri E. Nikonov, Ian A. Young
DAC12
2017 Technology Options for Beyond-CMOS
abstract
CMOS integrated circuit technology for computation is at an inflexion point. Although this is the technology which has enabled the semiconductor industry to make vast progress over the past 30-plus years, it is expected to see challenges going beyond the ten year horizon, particularly from an energy efficiency point of view. Thus it is extremely important for the semiconductor industry to discover a new integrated circuit technology which can carry us to the beyond CMOS era, so that the power-performance of computing can continue to improve. Currently, researchers are exploring novel device concepts and new information tokens as an alternative for CMOS technology. Examples of areas being actively researched are; quantum electronic devices, such as the tunneling field-effect transistor (TFET), and devices based on electron spin and nano-magnetics (spintronics). It is clear that choices will need to be made in the next 10 years to identify viable alternatives for CMOS by 2025. To prioritize and guide the research exploration in materials, devices and circuits, benchmarking methodology and metrics are being used. This talk will give an overview of the beyond CMOS device research horizon and the benchmarking of these devices for computation. A more detailed investigation of circuits based upon some promising beyond-CMOS devices will follow.
Ian A. Young
ISPD1
2013 Overview of Beyond-CMOS Devices and a Uniform Methodology for Their Benchmarking
abstract
Multiple logic devices are presently under study within the Nanoelectronic Research Initiative (NRI) to carry the development of integrated circuits beyond the complementary metal-oxide-semiconductor (CMOS) roadmap. Structure and operational principles of these devices are described. Theories used for benchmarking these devices are overviewed, and a general methodology is described for consistent estimates of the circuit area, switching time, and energy. The results of the comparison of the NRI logic devices using these benchmarks are presented.
Dmitri E. Nikonov, Ian A. Young
Proc. IEEE2
2007 A Blind Calibration Technique to Correct Memory Errors in Amplifier-sharing Pipelined ADCs
abstract
The authors present a statistics-based blind calibration technique for nonlinear memory errors in amplifier-sharing pipelined ADCs. The proposed method is fully digital and simple to implement. It detects memory errors in normal operation without system suspension for calibration. No special calibration signal or analog circuitry is necessary. Algorithm description and simulation results are presented. The proposed technique improves the power efficiency of pipelined ADCs by enabling sharing of low-gain operational amplifiers.
Munkyo Seo, Sopan Joshi, Ian A. Young
ISCAS3