Yuxuan Yin

dblp:287/5093 · DBLP profile ↗
← Back
11ranked-venue papers
8as first author
11since 2021 · last 2025
0009-0001-5504-3142ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 7 first-author · 9 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 3D Acceleration for Mixture-of-Experts and Multi-Head Attention Spiking Transformers with Dynamic Head Pruning
abstract
Spiking Neural Networks (SNNs) provide a brain-inspired and event-driven mechanism that is believed to be critical to unlock energy-efficient deep learning. On the other hand, mixture-of-experts (MoE) models mirror the parallel distributed processing of the nervous system, and expand model capacity without scaling up the number of computational operations. However, there is currently a lack of hardware support for highly parallel distributed processing in spiking based MoE models. This paper introduces the first 3D hardware architecture and design methodology for Mixture-of-Experts and Multi-Head Attention spiking transformers. By leveraging 3D integration with memory-on-logic and logic-on-logic stacking and exploring energy-efficient dynamic head pruning, we explore such brain-inspired accelerators with spatially stackable circuitry, demonstrating significant improvements of energy efficiency and latency over conventional 2D CMOS integration.
Boxun Xu, Junyoung Hwang, Pruek Vanna-Iampikul, Yuxuan Yin, Sung Kyu Lim, Peng Li 0001
ICCAD4
2025 Bishop: Sparsified Bundling Spiking Transformers on Heterogeneous Cores with Error-constrained Pruning
abstract
Spiking neural networks(SNNs) have emerged as a promising solution for deployment on resource-constrained edge devices and neuromorphic hardware due to their low power consumption.Spiking transformers, which integrate attention mechanisms similar to those found in artificial neural networks (ANNs), have recently exhibited impressive performance.However, these models are large in size and involve high-volume computation in both time and space, posing significant challenges for efficient hardware acceleration.We present Bishop, the first dedicated hardware accelerator architecture and HW/SW co-design framework for spiking transformers that optimally represents, manages, and processes spike-based workloads while exploring spatiotemporal sparsity and data reuse.Specifically, we introduce the concept of Token-Time Bundle (TTB), a container that bundles spiking data of a set of tokens over multiple time points.Our heterogeneous accelerator architecture Bishop concurrently processes workload packed in TTBs and explores intra-and inter-bundle multiple-bit weight reuse to significantly reduce memory access.Bishop utilizes a stratifier, a dense core array, and a sparse core array to process MLP blocks and projection layers.The stratifier routes high-density spiking activation workload to the dense core and low-density counterpart to the sparse core, ensuring optimized processing tailored to the given spatiotemporal sparsity level.To further reduce data access and computation, we introduce a novel Bundle Sparsity-Aware (BSA) training pipeline that enhances not only the overall but also structured TTB-level firing sparsity.Moreover, the processing efficiency of self-attention layers is boosted by the proposed Error-Constrained TTB Pruning (ECP), which trims activities in spiking queries, keys, and values both before and after the computation of spiking attention maps with a well-defined error bound.Finally, we design a reconfigurable TTB spiking attention core to efficiently compute spiking attention maps by executing highly simplified "AND" and "Accumulate" operations.On average, Bishop achieves a 5.91× speedup and 6.11×
Boxun Xu, Yuxuan Yin, Vikram Iyer, Peng Li 0001
ISCA2
2025 Transfer Learning for Minimum Operating Voltage Prediction in Advanced Technology Nodes: Leveraging Legacy Data and Silicon Odometer Sensing
abstract
Accurate prediction of chip performance is critical for ensuring energy efficiency and reliability in semiconductor manufacturing. However, developing minimum operating voltage (Vmin) prediction models at advanced technology nodes is challenging due to limited training data and the complex relationship between process variations and Vmin. To address these issues, we propose a novel transfer learning framework that leverages abundant legacy data from the 16nm technology node to enable accurate Vminprediction at the advanced 5nm node. A key innovation of our approach is the integration of input features derived from on-chip silicon odometer sensor data, which provide fine-grained characterization of localized process variations—an essential factor at the 5nm node—resulting in significantly improved prediction accuracy.
Yuxuan Yin, Rebecca Chen, Boxun Xu, Peng Li 0001
ITC1
2025 Data-Efficient Prediction of Minimum Operating Voltage via Inter- and Intra-Wafer Variation Alignment
abstract
Predicting the minimum operating voltage (Vmin) of chips stands as a crucial technique in enhancing the speed and reliability of manufacturing testing flow. However, existing Vminprediction methods often overlook various sources of variations in both training and deployment phases. Notably, overlooking wafer zone-to-zone (intra-wafer) variations and wafer-to-wafer (inter-wafer) variations diminishes the accuracy, data efficiency, and reliability of Vminpredictors. To address this challenge, we propose Restricted Bias Alignment (RBA), a novel data-efficient Vminprediction framework that introduces a variation alignment technique to simultaneously estimate inter- and intra-wafer variations. Furthermore, we propose utilizing class probe data to model inter-wafer variations for the first time.
Yuxuan Yin, Rebecca Chen, Peng Li 0001
VTS1
2025 Reliable Board-Level Degradation Prediction with Monotonic Segmented Regression under Noisy Measurement
abstract
The increasing complexity of electronic systems in autonomous electric vehicles necessitates robust methods for forecasting the degradation of critical components such as printed circuit boards (PCBs). Various time series forecasting methods have been investigated to predict in-situ resistance degradation under vibration loads. However, these methods failed to capture the degradation trend under strong measurement noise. This paper introduces Monotonic Segmented Linear Regression (MSLR), a novel approach designed to capture monotonic degradation trends in time series data under significant measurement noise. By incorporating monotonic constraints, MSLR effectively models the non-decreasing behavior characteristic of degradation processes. To further enhance reliability of the prediction, we integrate Adaptive Conformal Inference (ACI) with MSLR, enabling the estimation of statistically valid upper bounds for resistance degradation with high confidence. Extensive experiments demonstrate that MSLR outperforms state-of-the-art time series forecasting baselines on real-world PCB degradation datasets.
Yuxuan Yin, Rebecca Chen, Varun Thukral, Peng Li 0001
VTS1
2024 Semi-supervised Learning of Dynamical Systems with Neural Ordinary Differential Equations: A Teacher-Student Model Approach
abstract
Modeling dynamical systems is crucial for a wide range of tasks, but it remains challenging due to complex nonlinear dynamics, limited observations, or lack of prior knowledge. Recently, data-driven approaches such as Neural Ordinary Differential Equations (NODE) have shown promising results by leveraging the expressive power of neural networks to model unknown dynamics. However, these approaches often suffer from limited labeled training data, leading to poor generalization and suboptimal predictions. On the other hand, semi-supervised algorithms can utilize abundant unlabeled data and have demonstrated good performance in classification and regression tasks. We propose TS-NODE, the first semi-supervised approach to modeling dynamical systems with NODE. TS-NODE explores cheaply generated synthetic pseudo rollouts to broaden exploration in the state space and to tackle the challenges brought by lack of ground-truth system data under a teacher-student model. TS-NODE employs an unified optimization framework that corrects the teacher model based on the student's feedback while mitigating the potential false system dynamics present in pseudo rollouts. TS-NODE demonstrates significant performance improvements over a baseline Neural ODE model on multiple dynamical system modeling tasks.
Yu Wang 0167, Yuxuan Yin, Karthik Somayaji Nanjangud Suryanarayana, Ján Drgona, Malachi Schram, Mahantesh Halappanavar, Frank Liu 0001, Peng Li 0001
AAAI2
2024 Data-Efficient Conformalized Interval Prediction of Minimum Operating Voltage Capturing Process Variations
abstract
Accurate minimum operating voltage (Vmin) prediction is a critical element in manufacturing tests. Conventional methods lack coverage guarantees in interval predictions. Conformal Prediction (CP), a distribution-free machine learning approach, excels in providing rigorous coverage guarantees for interval predictions. However, standard CP predictors may fail due to a lack of knowledge of process variations. We address this challenge by providing principled conformalized interval prediction in the presence of process variations with high data efficiency, where the data from a few additional chips is utilized for calibration. We demonstrate the superiority of the proposed method on industrial 16nm chip data.
Yuxuan Yin, Rebecca Chen, Peng Li 0001
DAC1
2024 Reliable Interval Prediction of Minimum Operating Voltage Based on On-Chip Monitors via Conformalized Quantile Regression
abstract
Predicting the minimum operating voltage$V_{min}$of chips is one of the important techniques for improving the manufacturing testing flow, as well as ensuring the long-term reliability and safety of in-field systems. Current$V_{min}$prediction methods often provide only point estimates, necessitating additional techniques for constructing prediction confidence intervals to cover uncertainties caused by different sources of variations. While some existing techniques offer region predictions, but they rely on certain distributional assumptions and/or provide no coverage guarantees. In response to these limitations, we propose a novel distribution-free$V_{min}$interval estimation methodology possessing a theoretical guarantee of coverage. Our approach leverages conformalized quantile regression and on-chip monitors to generate reliable prediction intervals. We demonstrate the effectiveness of the proposed method on an industrial 5nm automotive chip dataset. Moreover, we show that the use of on-chip monitors can reduce the interval length significantly for$V_{min}$prediction.
Yuxuan Yin, Rebecca Chen, Peng Li 0001
DATE1
2024 ADO-LLM: Analog Design Bayesian Optimization with In-Context Learning of Large Language Models
abstract
Analog circuit design requires substantial human expertise and involvement, which is a significant roadblock to design productivity. Bayesian Optimization (BO), a popular machine-learning-based optimization strategy, has been leveraged to automate analog design given its applicability across various circuit topologies and technologies. Traditional BO methods employ black-box Gaussian Process surrogate models and optimized labeled data queries to find optimization solutions by trading off between exploration and exploitation. However, the search for the optimal design solution in BO can be expensive from both a computational and data usage point of view, particularly for high-dimensional optimization problems. This paper presents ADO-LLM, the first work integrating large language models (LLMs) with Bayesian Optimization for analog design optimization. ADO-LLM leverages the LLM's ability to infuse domain knowledge to rapidly generate viable design points to remedy BO's inefficiency in finding high-value design areas specifically under the limited design space coverage of the BO's probabilistic surrogate model. In the meantime, sampling of design points evaluated in the iterative BO process provides quality demonstrations for the LLM to generate high-quality design points while leveraging infused broad design knowledge. Furthermore, the diversity brought by BO's exploration enriches the contextual understanding of the LLM and allows it to more broadly search in the design space and prevent repetitive and redundant suggestions. We evaluate the proposed framework on two different types of analog circuits and demonstrate notable improvements in design efficiency and effectiveness.
Yuxuan Yin, Yu Wang 0167, Boxun Xu, Peng Li 0001
ICCAD1
2024 High-Dimensional Bayesian Optimization via Semi-Supervised Learning with Optimized Unlabeled Data Sampling
abstract
We introduce a novel semi-supervised learning approach, named Teacher-Student Bayesian Optimization ($\texttt{TSBO}$), integrating the teacher-student paradigm into BO to minimize expensive labeled data queries for the first time. $\texttt{TSBO}$ incorporates a teacher model, an unlabeled data sampler, and a student model. The student is trained on unlabeled data locations generated by the sampler, with pseudo labels predicted by the teacher. The interplay between these three components implements a unique selective regularization to the teacher in the form of student feedback. This scheme enables the teacher to predict high-quality pseudo labels, enhancing the generalization of the GP surrogate model in the search space. To fully exploit $\texttt{TSBO}$, we propose two optimized unlabeled data samplers to construct effective student feedback that well aligns with the objective of Bayesian optimization. Furthermore, we quantify and leverage the uncertainty of the teacher-student model for the provision of reliable feedback to the teacher in the presence of risky pseudo-label predictions. $\texttt{TSBO}$ demonstrates significantly improved sample-efficiency in several global optimization tasks under tight labeled data budgets. The implementation is available at https://github.com/reminiscenty/TSBO-Official.
Yuxuan Yin, Yu Wang 0167, Peng Li 0001
ICML1
2023 Domain-Specific Machine Learning Based Minimum Operating Voltage Prediction Using On-Chip Monitor Data
abstract
Determining the minimum operating voltage ($V_{min}$) of chip designs is critical for low power dissipation and assurance of quality and functional safety during manufacturing tests and in-field monitoring. We demonstrate how on-chip monitor data can be leveraged to provide accurate minimum operating voltage prediction using a domain-specific machine learning approach. Given limited measured chip data, the key challenge in developing a machine learning approach is to provide an accurate prediction while addressing overfitting and selecting a subset of optimal features. To this end, we propose to utilize a novel monotonic lattice neural network architecture that is geared towards accurate prediction by imposing domain-specific monotonic relationships between the input sensor data and$V_{min}$. Furthermore, we perform an effective feature selection by considering both the correlation between each feature and$V_{min}$as well as the co-linearity between the features. Experiments demonstrate superior performance in comparison with linear regression and conventional neural networks.
Yuxuan Yin, Rebecca Chen, Peng Li 0001
ITC1