Yanjing Li

dblp:62/201 · DBLP profile ↗
← Back
57ranked-venue papers
9as first author
37since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 32 · 4 first-author · 13 since 2021Artificial intelligence and machine learning · 22 · 5 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 10 since 2021Software engineering, systems software and programming languages · 7 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 UpDown: Efficient Manycore based on Many Threading and Scalable Memory Parallelism
abstract
Manycore architectures are a promising direction for single-chip performance. They typically use in-order cores and caches as building blocks and can produce good performance on regular applications. However, on irregular applications, they have low core utilization due to data-dependent control-flow and memory access.
Andronicus Rajasukumar, Ruiqi Xu 0001, Tianchi Zhang 0005, Yuqing Wang 0011, Tianshuo Su, Marziyeh Nourian, Jianru Ding, Jiya Su, Rajat Khandelwal, Alexander Fell, David F. Gleich, Yanjing Li, Henry Hoffmann, Andrew A. Chien
ICS12
2026 Associative Recurrent Bilinear Optimization for Domain-Generalized Binary Neural Networks
Sheng Xu 0007, Yanjing Li, Chuanjian Liu, Baochang Zhang 0001
Int. J. Comput. Vis.2
2026 UpDown: A Supercomputer Co-Designed for Scalable Graph Processing
abstract
Traditional supercomputers have focused on dense computation performance as exemplified by HPL. Graph processing applications differ with extreme irregularity ($10^{9}$imbalance in skewed, real-world graphs) that produces unpredictable work, parallelism, memory access, and communication. Together, these make scalable performance and programming difficult. We describe the UpDown system architecture, co-designed for irregular graph computations. UpDown provides efficient fine-grained thread invocations ($\sim$10 instructions), direct messaging (no network interface card) for scalable local and global messaging, and split-transaction memory operations that enable extremely high memory bandwidth. Combined with architectural support for global addressing and an aggressive network design, these UpDown features enable direct exploitation of edge and vertex parallelism, using it to deliver breakthrough graph processing performance and programmability. We evaluate the performance of the UpDown system using a challenging suite of graph applications (Pagerank, Breadth-first Search, Triangle Counting, Partial Match, etc). For a single-node, results show 100-fold performance advantage over multicore CPUs. Compared to today's fastest scalable parallel computers UpDown achieves 1000-fold performance increases. UpDown delivers these levels of performance with high-level programmability, these programs directly express vertex-edge parallelism which UpDown exploits directly in hardware.
Andrew A. Chien, Charles Colley, Jianru Ding, Alexander Fell, David F. Gleich, Henry Hoffmann, Moubarak Jeje, Rajat Khandelwal, Yanjing Li, Jose M. Monsalve Diaz, Marziyeh Nourian, Andronicus Rajasukumar, Jiya Su, Tianshuo Su, Yuqing Wang 0011, Ruiqi Xu 0001, Tianchi Zhang 0005
IEEE Trans. Parallel Distributed Syst.9
2025 SET: Spectral Enhancement for Tiny Object Detection
abstract
Deep learning has significantly advanced the object detection field. However, tiny object detection (TOD) remains a challenging problem. We provide a new analysis method to examine the TOD challenge through occlusion-based attribution analysis in the frequency domain. We observe that tiny objects become less distinct after feature encoding and can benefit from the removal of high-frequency information. In this paper, we propose a novel approach named Spectral Enhancement for Tiny object detection (SET), which amplifies the frequency signatures of tiny objects in a heterogeneous architecture. SET includes two modules. The Hierarchical Background Smoothing (HBS) module suppresses high-frequency noise in the background through adaptive smoothing operations. The Adversarial Perturbation Injection (API) module leverages adversarial perturbations to increase feature saliency in critical regions and prompt the refinement of object features during training. Extensive experiments on four datasets demonstrate the effectiveness of our method. Especially, SET boosts the prior art RFLA by 3.2% AP on the AI-TOD dataset.
Huixin Sun, Runqi Wang, Yanjing Li, Linlin Yang 0001, Shaohui Lin, Xianbin Cao 0001, Baochang Zhang 0001
CVPR3
2025 Uncertainty-Aware Gradient Stabilization for Small Object Detection
Huixin Sun, Yanjing Li, Linlin Yang 0001, Xianbin Cao 0001, Baochang Zhang 0001
ICCV2
2025 Efficient Low-Bit Quantization with Adaptive Scales for Multi-Task Co-Training
abstract
Co-training can achieve parameter-efficient multi-task models but remains unexplored for quantization-aware training. Our investigation shows that directly introducing co-training into existing quantization-aware training (QAT) methods results in significant performance degradation. Our experimental study identifies that the primary issue with existing QAT methods stems from the inadequate activation quantization scales for the co-training framework. To address this issue, we propose Task-Specific Scales Quantization for Multi-Task Co-Training (TSQ-MTC) to tackle mismatched quantization scales. Specifically, a task-specific learnable multi-scale activation quantizer (TLMAQ) is incorporated to enrich the representational ability of shared features for different tasks. Additionally, we find that in the deeper layers of the Transformer model, the quantized network suffers from information distortion within the attention quantizer. A structure-based layer-by-layer distillation (SLLD) is then introduced to ensure that the quantized features effectively preserve the information from their full-precision counterparts. Our extensive experiments in two co-training scenarios demonstrate the effectiveness and versatility of TSQ-MTC. In particular, we successfully achieve a 4-bit quantized low-level visual foundation model based on IPT, which attains a PSNR comparable to the full-precision model while offering a $7.99\times$ compression ratio in the $\times4$ super-resolution task on the Set5 benchmark.
Linlin Yang 0001, Yanjing Li, Guodong Guo, Xianbin Cao 0001, Baochang Zhang 0001
ICLR4
2025 SWIPER: Minimizing Fault-Tolerant Quantum Program Latency via Speculative Window Decoding
abstract
Real-time decoding is a key ingredient in future fault-tolerant quantum systems, yet many decoders are too slow to run in real time.Prior work has shown that parallel window decoding can scalably meet throughput requirements in the presence of increasing decoding times.However, windowed decoding require that some decoding tasks be delayed until others have completed, which can be problematic during time-sensitive operations such as T gate teleportation, leading to suboptimal program runtimes.To alleviate this, we introduce SWIPER, a speculative window decoder.Taking inspiration from branch prediction in classical computer architecture, SWIPER utilizes a light-weight speculation step to predict data dependencies between adjacent decoding windows, allowing multiple layers of decoding tasks to be resolved simultaneously.Through a state-of-the-art compilation pipeline and a detailed open-source simulator, we find that SWIPER reduces application runtimes by 40% on average compared to prior parallel window decoders.
Joshua Viszlai, Jason Chadwick, Gokul Subramanian Ravi, Yanjing Li, Fred Chong
ISCA5
2025 Silent Data Corruption: Advancing Detection, Diagnosis, and Mitigation Strategies
abstract
Silent Data Corruptions (SDCs) pose a critical challenge to computer system reliability, arising from vulnerabilities across different layers of the computing stack. This paper addresses this challenge through three complementary contributions that systematically target SDCs from hardware manufacturing to application-level resilience. First, we analyze timing failures caused by random process variations in advanced technology nodes, revealing that extreme slow paths at lower voltages are dominated by single weak transistors—insights crucial for manufacturing and in-field testing. Second, we introduce an LLM-driven framework that generates targeted functional test programs to induce SDCs, demonstrating its effectiveness in stressing hardware, uncovering latent vulnerabilities, and increasing energy consumption in a given device under test (DUT), making it a valuable tool for in-field testing. Third, as machine learning continues to drive advancements across critical domains such as healthcare, finance, and autonomous systems, ensuring its reliability is paramount. However, the susceptibility of these applications to SDCs threatens their reliability and robustness. To address this, we propose Fidelity-Q, a novel fault injection methodology to evaluate the impact of SDCs on Quantized Neural Networks (QNNs), showing that lower-bit quantization increases error susceptibility. Collectively, these contributions provide a comprehensive approach to identifying, analyzing, and mitigating SDCs across the computing stack, from hardware testing to machine learning applications.
Peter Domanski, Mukarram Ali Faridi, Gabriel Kaunang, Wilson Pradeep, Adit D. Singh, Alfian Amrizal, Yanjing Li, Farshad Firouzi, Krishnendu Chakrabarty
VTS7
2025 Learning Accurate Low-bit Quantization towards Efficient Computational Imaging
Sheng Xu 0007, Yanjing Li, Chuanjian Liu, Baochang Zhang 0001
Int. J. Comput. Vis.2
2025 M3DP: Optimizing 2D vision tasks with minimal 3D object information
Yanjing Li, Linlin Yang 0001, Xinkai Liang, Xianbin Cao 0001, Qi Wang 0009, Baochang Zhang 0001
Neurocomputing2
2025 Calibrated gradient descent of convolutional neural networks for embodied visual recognition
Sheng Xu 0007, Lian Zhuo, Baochang Zhang 0001, Yanjing Li, Guodong Guo
Image Vis. Comput.5
2024 Bi-ViT: Pushing the Limit of Vision Transformer Quantization
abstract
Vision transformers (ViTs) quantization offers a promising prospect to facilitate deploying large pre-trained networks on resource-limited devices. Fully-binarized ViTs (Bi-ViT) that pushes the quantization of ViTs to its limit remain largely unexplored and a very challenging task yet, due to their unacceptable performance. Through extensive empirical analyses, we identify the severe drop in ViT binarization is caused by attention distortion in self-attention, which technically stems from the gradient vanishing and ranking disorder. To address these issues, we first introduce a learnable scaling factor to reactivate the vanished gradients and illustrate its effectiveness through theoretical and experimental analyses. We then propose a ranking-aware distillation method to rectify the disordered ranking in a teacher-student framework. Bi-ViT achieves significant improvements over popular DeiT and Swin backbones in terms of Top-1 accuracy and FLOPs. For example, with DeiT-Tiny and Swin-Tiny, our method significantly outperforms baselines by 22.1% and 21.4% respectively, while 61.5x and 56.1x theoretical acceleration in terms of FLOPs compared with real-valued counterparts on ImageNet. Our codes and models are attached on https://github.com/YanjingLi0202/Bi-ViT/ .
Yanjing Li, Sheng Xu 0007, Mingbao Lin, Xianbin Cao 0001, Chuanjian Liu, Baochang Zhang 0001
AAAI1
2024 YFlows: Systematic Dataflow Exploration and Code Generation for Efficient Neural Network Inference using SIMD Architectures on CPUs
abstract
We address the challenges associated with deploying neural networks on CPUs, with a particular focus on minimizing inference time while maintaining accuracy. Our novel approach is to use the dataflow (i.e., computation order) of a neural network to explore data reuse opportunities using heuristic-guided analysis and a code generation framework, which enables exploration of various Single Instruction, Multiple Data (SIMD) implementations to achieve optimized neural network execution. Our results demonstrate that the dataflow that keeps outputs in SIMD registers while also maximizing both input and weight reuse consistently yields the best performance for a wide variety of inference workloads, achieving up to 3x speedup for 8-bit neural networks, and up to 4.8x speedup for binary neural networks, respectively, over the optimized implementations of neural networks today.
Cyrus Zhou, Zack Hassman, Dhirpal Shah, Vaughn Richard, Yanjing Li
CC5
2024 Learning 1-Bit Tiny Object Detector with Discriminative Feature Refinement
abstract
1-bit detectors show impressive performance comparable to their real-valued counterparts when detecting commonly sized objects while exhibiting significant performance degradation on tiny objects. The challenge stems from the fact that high-level features extracted by 1-bit convolutions seem less compelling to reveal the discriminative foreground features. To address these issues, we introduce a Discriminative Feature Refinement method for 1-bit Detectors (DFR-Det), aiming to enhance the discriminative ability of foreground representation for tiny objects in aerial images. This is accomplished by refining the feature representation using an information bottleneck (IB) to achieve a distinctive representation of tiny objects. Specifically, we introduce a new decoder with a foreground mask, aiming to enhance the discriminative ability of high-level features for the target but suppress the background impact. Additionally, our decoder is simple but effective and can be easily mounted on existing detectors without extra burden added to the inference procedure. Extensive experiments on various tiny object detection (TOD) tasks demonstrate DFR-Det’s superiority over state-of-the-art 1-bit detectors. For example, 1-bit FCOS achieved by DFR-Det achieves the 12.8% AP on AI-TOD dataset, approaching the performance of the real-valued counterpart.
Sheng Xu 0007, Yanjing Li, Mingbao Lin, Baochang Zhang 0001, David S. Doermann
ICML3
2024 Drop-Connect as a Fault-Tolerance Approach for RRAM-based Deep Neural Network Accelerators
abstract
Resistive random-access memory (RRAM) is widely recognized as a promising emerging hardware platform for deep neural networks (DNNs). Yet, due to manufacturing limitations, current RRAM devices are highly susceptible to hardware defects, which poses a significant challenge to their practical applicability. In this paper, we present a machine learning technique that enables the deployment of defect-prone RRAM accelerators for DNN applications, without necessitating modifying the hardware, retraining of the neural network, or implementing additional detection circuitry/logic. The key idea involves incorporating a drop-connect inspired approach during the training phase of a DNN, where random subsets of weights are selected to emulate fault effects (e.g., set to zero to mimic stuck-at-1 faults), thereby equipping the DNN with the ability to learn and adapt to RRAM defects with the corresponding fault rates. Our results demonstrate the viability of the dropconnect approach, coupled with various algorithm and system-level design and trade-off considerations. We show that, even in the presence of high defect rates (e.g., up to 30%), the degradation of DNN accuracy can be as low as less than 1% compared to that of the fault-free version, while incurring minimal system-level runtime/energy costs.
Mingyuan Xiang, Xuhan Xie, Pedro Savarese, Michael Maire, Yanjing Li
VTS6
2023 Resilient Binary Neural Network
abstract
Binary neural networks (BNNs) have received ever-increasing popularity for their great capability of reducing storage burden as well as quickening inference time. However, there is a severe performance drop compared with {real-valued} networks, due to its intrinsic frequent weight oscillation during training. In this paper, we introduce a Resilient Binary Neural Network (ReBNN) to mitigate the frequent oscillation for better BNNs' training. We identify that the weight oscillation mainly stems from the non-parametric scaling factor. To address this issue, we propose to parameterize the scaling factor and introduce a weighted reconstruction loss to build an adaptive training objective. For the first time, we show that the weight oscillation is controlled by the balanced parameter attached to the reconstruction loss, which provides a theoretical foundation to parameterize it in back propagation. Based on this, we learn our ReBNN by calculating the balanced parameter based on its maximum magnitude, which can effectively mitigate the weight oscillation with a resilient training process. Extensive experiments are conducted upon various network models, such as ResNet and Faster-RCNN for computer vision, as well as BERT for natural language processing. The results demonstrate the overwhelming performance of our ReBNN over prior arts. For example, our ReBNN achieves 66.9% Top-1 accuracy with ResNet-18 backbone on the ImageNet dataset, surpassing existing state-of-the-arts by a significant margin. Our code is open-sourced at https://github.com/SteveTsui/ReBNN.
Sheng Xu 0007, Yanjing Li, Teli Ma, Mingbao Lin, Hao Dong 0003, Baochang Zhang 0001, Peng Gao 0007, Jinhu Lü 0001
AAAI2
2023 Implicit Diffusion Models for Continuous Super-Resolution
abstract
Image super-resolution (SR) has attracted increasing attention due to its widespread applications. However, current SR methods generally suffer from over-smoothing and artifacts, and most work only with fixed magnifications. This paper introduces an Implicit Diffusion Model (IDM) for high-fidelity continuous image super-resolution. IDM integrates an implicit neural representation and a denoising diffusion model in a unified end-to-end framework, where the implicit neural representation is adopted in the decoding process to learn continuous-resolution representation. Furthermore, we design a scale-adaptive conditioning mechanism that consists of a low-resolution (LR) conditioning network and a scaling factor. The scaling factor regulates the resolution and accordingly modulates the proportion of the LR information and generated features in the final output, which enables the model to accommodate the continuous-resolution requirement. Extensive experiments validate the effectiveness of our IDM and demonstrate its superior performance over prior arts. The source code will be available at https://github.com/Ree1s/IDM.
Sicheng Gao, Xuhui Liu, Bohan Zeng, Sheng Xu 0007, Yanjing Li, Xiaoyan Luo, Jianzhuang Liu, Xiantong Zhen, Baochang Zhang 0001
CVPR5
2023 Q-DETR: An Efficient Low-Bit Quantized Detection Transformer
abstract
The recent detection transformer (DETR) has advanced object detection, but its application on resource-constrained devices requires massive computation and memory resources. Quantization stands out as a solution by representing the network in low-bit parameters and operations. However, there is a significant performance drop when performing low-bit quantized DETR (Q-DETR) with existing quantization methods. We find that the bottle-necks of Q-DETR come from the query information distortion through our empirical analyses. This paper addresses this problem based on a distribution rectification distillation (DRD). We formulate our DRD as a bi-level optimization problem, which can be derived by generalizing the information bottleneck (IB) principle to the learning of Q-DETR. At the inner level, we conduct a distribution alignment for the queries to maximize the self-information entropy. At the upper level, we introduce a new foreground-aware query matching scheme to effectively transfer the teacher information to distillation-desired features to minimize the conditional information entropy. Extensive experimental results show that our method performs much better than prior arts. For example, the 4-bit Q-DETR can theoretically accelerate DETR with ResNet-50 backbone by 6.6× and achieve 39.4% AP, with only 2.6% performance gaps than its real-valued counterpart on the COCO dataset11Code: https://github.com/SteveTsui/Q-DETR.
Sheng Xu 0007, Yanjing Li, Mingbao Lin, Peng Gao 0007, Guodong Guo, Jinhu Lü 0001, Baochang Zhang 0001
CVPR2
2023 Memristor-Spikelearn: A Spiking Neural Network Simulator for Studying Synaptic Plasticity under Realistic Device and Circuit Behaviors
abstract
We present the Memristor-Spikelearn simulator (open-sourced), which is capable of incorporating detailed mem-ristor and circuit models in simulation to enable thorough study of synaptic plasticity in spiking neural networks under realistic device and circuit behaviors. Using this simulator, we demonstrate that: (1) a detailed device model is essential for simulating synaptic plasticity workloads, because results obtained using a simplified model can be misleading (e.g., it can overestimate test accuracy by up to 21.9%); (2) detailed simulation helps to determine the proper range of conductance values to represent weights, which is critical in order to achieve the desired accuracy -energy tradeoff (e.g., increasing the conductance values by$10\times$can increase accuracy from 70% to 83% at the price of$20\times$higher energy); and (3) detailed simulation also helps to determine an optimized circuit structure, which is another important design parameter that can yield different accuracy -energy tradeoffs.
Angel Yanguas-Gil, Sandeep Madireddy, Yanjing Li
DATE4
2023 Understanding Permanent Hardware Failures in Deep Learning Training Accelerator Systems
abstract
Hardware failures pose critical threats to deep neural network (DNN) training workloads, and the urgency of tackling this challenge (known as the Silent Data Corruption challenge in a broader context) has been raised widely by the industry. Based on industry reports, a large number of the failures observed in real systems are permanent hardware failures in logic. However, there is a very limited understanding of the effects that these failures can impose on DNN training workloads. In this paper, we present the first resilience study on this subject, focusing on deep learning (DL) training accelerator systems. We developed a fault injection framework to accurately simulate the effects of permanent faults, and conducted 100K fault injection experiments. Our results provide the fundamental understanding on how logic permanent hardware failures affect training workloads and eventually generate unexpected training outcomes. Based on this new knowledge, we developed efficient software-based detection and recovery techniques to mitigate logic permanent hardware failures that are likely to generate unexpected outcomes. Evaluation on Google Cloud TPUs shows that our techniques are effective and practical: they require 15−25 lines of code change, and introduce 0.004%−0.025% performance/energy overhead for various representative neural network models.
Yi He 0010, Yanjing Li
ETS2
2023 Representation Disparity-aware Distillation for 3D Object Detection
abstract
In this paper, we focus on developing knowledge distillation (KD) for compact 3D detectors. We observe that off-the-shelf KD methods manifest their efficacy only when the teacher model and student counterpart share similar intermediate feature representations. This might explain why they are less effective in building extreme-compact 3D detectors where significant representation disparity arises due primarily to the intrinsic sparsity and irregularity in 3D point clouds. This paper presents a novel representation disparity-aware distillation (RDD) method to address the representation disparity issue and reduce performance gap between compact students and over-parameterized teachers. This is accomplished by building our RDD from an innovative perspective of information bottleneck (IB), which can effectively minimize the disparity of proposal region pairs from student and teacher in features and logits. Extensive experiments are performed to demonstrate the superiority of our RDD over existing KD methods. For example, our RDD increases mAP of CP-Voxel-S to 57.1% on nuScenes dataset, which even surpasses teacher performance while taking up only 42% FLOPs.
Yanjing Li, Sheng Xu 0007, Mingbao Lin, Jihao Yin, Baochang Zhang 0001, Xianbin Cao 0001
ICCV1
2023 Understanding and Mitigating Hardware Failures in Deep Learning Training Systems
abstract
Deep neural network (DNN) training workloads are increasingly susceptible to hardware failures in datacenters. For example, Google experienced "mysterious, difficult to identify problems" in their TPU training systems due to hardware failures [7]. Although these particular problems were subsequently corrected through significant efforts, they have raised the urgency of addressing the growing challenges emerging from hardware failures impacting many DNN training workloads.
Yi He 0010, Mike Hutton, Robert De Gruijl, Rama Govindaraju, Nishant Patil, Yanjing Li
ISCA7
2023 Q-DM: An Efficient Low-bit Quantized Diffusion Model
abstract
Denoising diffusion generative models are capable of generating high-quality data, but suffers from the computation-costly generation process, due to a iterative noise estimation using full-precision networks. As an intuitive solution, quantization can significantly reduce the computational and memory consumption by low-bit parameters and operations. However, low-bit noise estimation networks in diffusion models (DMs) remain unexplored yet and perform much worse than the full-precision counterparts as observed in our experimental studies. In this paper, we first identify that the bottlenecks of low-bit quantized DMs come from a large distribution oscillation on activations and accumulated quantization error caused by the multi-step denoising process. To address these issues, we first develop a Timestep-aware Quantization (TaQ) method and a Noise-estimating Mimicking (NeM) scheme for low-bit quantized DMs (Q-DM) to effectively eliminate such oscillation and accumulated error respectively, leading to well-performed low-bit DMs. In this way, we propose an efficient Q-DM to calculate low-bit DMs by considering both training and inference process in the same framework. We evaluate our methods on popular DDPM and DDIM models. Extensive experimental results show that our method achieves a much better performance than the prior arts. For example, the 4-bit Q-DM theoretically accelerates the 1000-step DDPM by 7.8x and achieves a FID score of 5.17, on the unconditional CIFAR-10 dataset.
Yanjing Li, Sheng Xu 0007, Xianbin Cao 0001, Baochang Zhang 0001
NeurIPS1
2023 Symmetry-Informed Geometric Representation for Molecules, Proteins, and Crystalline Materials
abstract
Artificial intelligence for scientific discovery has recently generated significant interest within the machine learning and scientific communities, particularly in the domains of chemistry, biology, and material discovery. For these scientific problems, molecules serve as the fundamental building blocks, and machine learning has emerged as a highly effective and powerful tool for modeling their geometric structures. Nevertheless, due to the rapidly evolving process of the field and the knowledge gap between science ({\eg}, physics, chemistry, & biology) and machine learning communities, a benchmarking study on geometrical representation for such data has not been conducted. To address such an issue, in this paper, we first provide a unified view of the current symmetry-informed geometric methods, classifying them into three main categories: invariance, equivariance with spherical frame basis, and equivariance with vector frame basis. Then we propose a platform, coined Geom3D, which enables benchmarking the effectiveness of geometric strategies. Geom3D contains 16 advanced symmetry-informed geometric representation models and 14 geometric pretraining methods over 52 diverse tasks, including small molecules, proteins, and crystalline materials. We hope that Geom3D can, on the one hand, eliminate barriers for machine learning researchers interested in exploring scientific problems; and, on the other hand, provide valuable guidance for researchers in computational chemistry, structural biology, and materials science, aiding in the informed selection of representation techniques for specific applications. The source code is available on \href{https://github.com/chao1224/Geom3D}{the GitHub repository}.
Shengchao Liu, Weitao Du, Yanjing Li, Zhuoxinran Li, Zhiling Zheng, Chenru Duan, Zhiming Ma, Omar Yaghi, Anima Anandkumar, Christian Borgs, Jennifer T. Chayes, Jian Tang 0005
NeurIPS3
2023 DCP-NAS: Discrepant Child-Parent Neural Architecture Search for 1-bit CNNs
Yanjing Li, Sheng Xu 0007, Xianbin Cao 0001, Lian Zhuo, Baochang Zhang 0001, Tian Wang 0002, Guodong Guo
Int. J. Comput. Vis.1
2023 A Hybrid Optical-Electrical Analog Deep Learning Accelerator Using Incoherent Optical Signals
abstract
Optical deep learning (DL) accelerators have attracted significant interests due to their latency and power advantages. In this article, we focus on incoherent optical designs. A significant challenge is that there is no known solution to perform single-wavelength accumulation (a key operation required for DL workloads) using incoherent optical signals efficiently. Therefore, we devise a hybrid approach, where accumulation is done in the electrical domain, and multiplication is performed in the optical domain. The key technology enabler of our design is the transistor laser, which performs electrical-to-optical and optical-to-electrical conversions efficiently. Through detailed design and evaluation of our design, along with a comprehensive benchmarking study against state-of-the-art RRAM-based designs, we derive the following key results: (1) For a four-layer multilayer perceptron network, our design achieves 115× and 17.11× improvements in latency and energy, respectively, compared to the RRAM-based design. We can take full advantage of the speed and energy benefits of the optical technology because the inference task can be entirely mapped onto our design. (2) For a complex workload (Resnet50), weight reprogramming is needed, and intermediate results need to be stored/re-fetched to/from memories. In this case, for the same area, our design still outperforms the RRAM-based design by 15.92× in inference latency, and 8.99× in energy.
Mingdai Yang, Qiuwen Lou, Ramin Rajaei, Mohammad Reza Jokar, Junyi Qiu, Aditi Udupa, Fred Chong, John M. Dallesasse, Milton Feng, Lynford L. Goddard, Xiaobo Sharon Hu, Yanjing Li
ACM J. Emerg. Technol. Comput. Syst.13
2022 Recurrent Bilinear Optimization for Binary Neural Networks
Sheng Xu 0007, Yanjing Li, Teli Ma, Baochang Zhang 0001, Peng Gao 0007, Yu Qiao 0001, Jinhu Lü 0001, Guodong Guo
ECCV (24)2
2022 IDa-Det: An Information Discrepancy-Aware Distillation for 1-Bit Detectors
Sheng Xu 0007, Yanjing Li, Bohan Zeng, Teli Ma, Baochang Zhang 0001, Xianbin Cao 0001, Peng Gao 0007, Jinhu Lü 0001
ECCV (11)2
2022 Achieving Automotive Safety Requirements through Functional In-Field Self-Test for Deep Learning Accelerators
abstract
Deep learning (DL) accelerators are prominent in automotive systems, and it is essential to guarantee that these accelerators can meet the stringent automotive safety standard even in the presence of various hardware failures. In our previous work [1], we developed an efficient functional in-field self-test generation technique targeting DL accelerators, which achieves high (99.9%) stuck-at fault coverage. In this paper, we present an industry case study that extends our previous work to generate functional in-field self-tests with high transition test coverage, which is critical for screening timing degradation (e.g., caused by circuit aging). We will first present an overview of the general in-vehicle system architecture and discuss reliability/safety requirements and goals. Next, we will discuss the details of our functional in-field self-test generation technique for the transition fault model. Finally, through detailed evaluation on an industrial DL accelerator design, we will show that our approach is able to achieve extremely high transition test coverage (> 99.0%), thereby successfully achieving the reliability/safety requirements for DL accelerators in automotive applications. Moreover, the total in-field self-test time and test storage costs of our technique are low, within the required constraints.
Takumi Uezono, Yi He 0010, Yanjing Li
ITC3
2022 Q-ViT: Accurate and Fully Quantized Low-bit Vision Transformer
abstract
The large pre-trained vision transformers (ViTs) have demonstrated remarkable performance on various visual tasks, but suffer from expensive computational and memory cost problems when deployed on resource-constrained devices. Among the powerful compression approaches, quantization extremely reduces the computation and memory consumption by low-bit parameters and bit-wise operations. However, low-bit ViTs remain largely unexplored and usually suffer from a significant performance drop compared with the real-valued counterparts. In this work, through extensive empirical analysis, we first identify the bottleneck for severe performance drop comes from the information distortion of the low-bit quantized self-attention map. We then develop an information rectification module (IRM) and a distribution guided distillation (DGD) scheme for fully quantized vision transformers (Q-ViT) to effectively eliminate such distortion, leading to a fully quantized ViTs. We evaluate our methods on popular DeiT and Swin backbones. Extensive experimental results show that our method achieves a much better performance than the prior arts. For example, our Q-ViT can theoretically accelerates the ViT-S by 6.14x and achieves about 80.9% Top-1 accuracy, even surpassing the full-precision counterpart by 1.0% on ImageNet dataset. Our codes and models are attached on https://github.com/YanjingLi0202/Q-ViT
Yanjing Li, Sheng Xu 0007, Baochang Zhang 0001, Xianbin Cao 0001, Peng Gao 0007, Guodong Guo
NeurIPS1
2022 Not All Bits have Equal Value: Heterogeneous Precisions via Trainable Noise
abstract
We study the problem of training deep networks while quantizing parameters and activations into low-precision numeric representations, a setting central to reducing energy consumption and inference time of deployed models. We propose a method that learns different precisions, as measured by bits in numeric representations, for different weights in a neural network, yielding a heterogeneous allocation of bits across parameters. Learning precisions occurs alongside learning weight values, using a strategy derived from a novel framework wherein the intractability of optimizing discrete precisions is approximated by training per-parameter noise magnitudes. We broaden this framework to also encompass learning precisions for hidden state activations, simultaneously with weight precisions and values. Our approach exposes the objective of constructing a low-precision inference-efficient model to the entirety of the training process. Experiments show that it finds highly heterogeneous precision assignments for CNNs trained on CIFAR and ImageNet, improving upon previous state-of-the-art quantization methods. Our improvements extend to the challenging scenario of learning reduced-precision GANs.
Pedro Savarese, Yanjing Li, Michael Maire
NeurIPS3
2022 Special Session: On the Reliability of Conventional and Quantum Neural Network Hardware
abstract
Neural Networks (NNs) are being extensively used in critical applications such as aerospace, healthcare, autonomous driving, and military, to name a few. Limited precision of the underlying hardware platforms, permanent and transient faults injected unintentionally as well as maliciously, and voltage/temperature fluctuations can potentially result in malfunctions in NNs with consequences ranging from substantial reduction in the network accuracy to jeopardizing the correct prediction of the network in worst cases. To alleviate such reliability concerns, this paper discusses the state-of-the-art reliability enhancement schemes that can be tailored for deep learning accelerators. We will discuss the errors associated with the hardware implementation of Deep-Learning (DL) algorithms along with their corresponding countermeasures. An in-field self-test methodology with a high test coverage is introduced, and an accurate high-level framework, so-called FIdelity, is proposed that enables the designers to evaluate DL accelerators in presence of such errors. Then, a state-of-the-art robustness-preserving training algorithm based on the Hessian Regularization is introduced. This algorithm alleviates the perturbations during inference time with negligible degradation in the accuracy of the network. Finally, Quantum Neural Networks (QNNs) and the methods to make them resilient against a variety of vulnerabilities such as fault injection, spatial and temporal variations in Qubits, and noise in QNNs are discussed.
Mehdi Sadi, Yi He 0010, Yanjing Li, Mahabubul Alam, Satwik Kundu, Swaroop Ghosh, Javad Bahrami, Naghmeh Karimi
VTS3
2022 Filter pruning via expectation-maximization
Sheng Xu 0007, Yanjing Li, Linlin Yang 0001, Baochang Zhang 0001, Dianmin Sun
Neural Comput. Appl.2
2021 POEM: 1-bit Point-wise Operations based on Expectation-Maximization for Efficient Point Cloud Processing
Sheng Xu 0007, Junhe Zhao, Yanjing Li, Baochang Zhang 0001, Guodong Guo
BMVC3
2021 A Hybrid Optical-Electrical Analog Deep Learning Accelerator Using Incoherent Optical Signals
abstract
We present a hybrid optical-electrical analog deep learning (DL) accelerator, the first work to use incoherent optical signals for DL workloads. Incoherent optical designs are more attractive than coherent ones as the former can be more easily realized in practice. However, a significant challenge in analog DL accelerators, where multiply-accumulate operations are dominant, is that there is no known solution to perform accumulation using incoherent optical signals. We overcome this challenge by devising a hybrid approach: accumulation is done in the electrical domain, while multiplication is performed in the optical domain. The key technology enabler of our design is the transistor laser, which performs electrical-to-optical and optical-to-electrical conversions efficiently to tightly integrate electrical and optical devices into compact circuits. As such, our design fully realizes the ultra high-speed and high-energy-efficiency advantages of analog and optical computing. Our evaluation results using the MNIST benchmark show that our design achieves 2214× and 65× improvements in latency and energy, respectively, compared to a state-of-the-art memristor-based analog design.
Mingdai Yang, Mohammad Reza Jokar, Junyi Qiu, Qiuwen Lou, Aditi Udupa, Fred Chong, John M. Dallesasse, Milton Feng, Lynford L. Goddard, Xiaobo Sharon Hu, Yanjing Li
ACM Great Lakes Symposium on VLSI12
2021 Efficient Functional In-Field Self-Test for Deep Learning Accelerators
abstract
We present a technique that generates high-quality functional in-field self-tests specifically targeting deep learning (DL) accelerators. These functional tests can be applied in the field during normal operation of a DL accelerator, which is crucial to ensure that the safety and/or reliability requirements are met for any given application, including safety-critical applications such as self-driving cars, robotics, and more.Our technique takes advantage of special architectural characteristics and application properties to achieve high functional test coverage while incurring minimal system-level costs. Moreover, we devise different strategies for the compute units (which support computation operations) and the control units (which control data movement) because these two types of units exhibit different properties. For the compute units of a DL accelerator, we first use combinational ATPG to generate test patterns with high test coverage, which is possible because these units do not contain complex sequential logic. Next, we map the ATPG patterns to one or more equivalent deep neural networks (DNNs) that can be directly executed on the accelerator, which is possible given the well-defined dataflow/reuse algorithm of a DL accelerator. For the control units, we leverage the property that typically only one or a few fixed DNNs are deployed at a time in many application domains (e.g., self-driving cars). Thus, it is sufficient to target only the faults that can directly affect the correctness of the DNNs that are currently deployed. This is done by executing different layers of each target DNN using carefully-crafted input and weight values to maximize test coverage while minimizing test time.We apply our technique using Nvidia’s open-source accelerator as a case study to demonstrate its efficacy. Our results show that our technique achieves high test coverage. For the compute units, 99.9% single stuck-at functional test coverage is achieved. For the control units, we are able to prove that, given any target DNN, 100% coverage can be achieved for a large class of single and multiple fault models. The in-field functional self-test time is also very low, < 17 ms for various representative DNNs. These functional tests can be applied during boot-up, reset, and even concurrently with normal operation by executing DNN test programs directly on the accelerator, without requiring any test support in the hardware.
Yi He 0010, Takumi Uezono, Yanjing Li
ITC3
2021 Tfcancer: a manually curated database of transcription factors associated with human cancers
abstract
SUMMARY: Transcription factors (TFs) are critical regulation elements and its dysregulation can lead to a variety of cancers. However, currently, there are no such online resources for large-scale collection, storage and analysis of TF-cancer associations in those cancers. To fill this gap, we present a database called TFcancer (http://lcbb.swjtu.edu.cn/tfcancer/), which contains 3136 experimentally supported associations between 364 TFs and 33 TCGA cancers by manually curating more than 1800 literature. TFcancer mainly concentrates on four aspects: TF expression, molecular alteration, regulatory relationships between TFs and target genes, and biological processes and signaling pathways of TFs in cancers. TFcancer not only provides a user-friendly interface for browsing and searching but also allows flexible data downloading and user data submitting. It is believed that TFcancer is a helpful and valuable resource for researchers who seek to understand the functions and molecular mechanisms of TFs involved in human cancers. AVAILABILITY AND IMPLEMENTATION: The TFcancer are freely available at http://lcbb.swjtu.edu.cn/tfcancer/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zhengtang Tan, Yanjing Li, Wenzhu Wang, Mei Lang, Changying Li, Zhiyun Guo
Bioinform.3
2020 Baldur: A Power-Efficient and Scalable Network Using All-Optical Switches
abstract
We present the first all-optical network, Baldur, to enable power-efficient and high-speed communications in future exascale computing systems. The essence of Baldur is its ability to perform packet routing on-the-fly in the optical domain using an emerging technology called the transistor laser (TL), which presents interesting opportunities and challenges at the system level. Optical packet switching readily eliminates many inefficiencies associated with the crossings between optical and electrical domains. However, TL gates consume high power at the current technology node, which makes TL-based buffering and optical clock recovery impractical. Consequently, we must adopt novel (bufferless and clock-less) architecture and design approaches that are substantially different from those used in current networks. At the architecture level, we support a bufferless design by turning to techniques that have fallen out of favor for current networks. Baldur uses a low-radix, multi-stage network with a simple routing algorithm that drops packets to handle congestion, and we further incorporate path multiplicity and randomness to minimize packet drops. This design also minimizes the number of TL gates needed in each switch. At the logic design level, a non-conventional, length-based data encoding scheme is used to eliminate the need for clock recovery. We thoroughly validate and evaluate Baldur using a circuit simulator and a network simulator. Our results show that Baldur achieves up to 3,000X lower average latency while consuming 3.2X-26.4X less power than various state-of-the art networks under a wide variety of traffic patterns and real workloads, for the scale of 1,024 server nodes. Baldur is also highly scalable, since its power per node stays relatively constant as we increase the network size to over 1 million server nodes, which corresponds to 14.6X-31.0X power improvements compared to state-of-the-art networks at this scale.
Mohammad Reza Jokar, Junyi Qiu, Fred Chong, Lynford L. Goddard, John M. Dallesasse, Milton Feng, Yanjing Li
HPCA7
2020 FIdelity: Efficient Resilience Analysis Framework for Deep Learning Accelerators
abstract
We present a resilience analysis framework, called FIdelity, to accurately and quickly analyze the behavior of hardware errors in deep learning accelerators. Our framework enables resilience analysis starting from the very beginning of the design process to ensure that the reliability requirements are met, so that these accelerators can be safely deployed for a wide range of applications, including safety-critical applications such as self-driving cars.Existing resilience analysis techniques suffer from the following limitations: 1. general-purpose hardware techniques can achieve accurate results, but they require access to RTL to perform time-consuming RTL simulations, which is not feasible for early design exploration; 2. general-purpose software techniques can produce results quickly, but they are highly inaccurate; 3. techniques targeting deep learning accelerators only focus on memory errors.Our FIdelity framework overcomes these limitations. FIdelity only requires a minimal amount of high-level design information that can be obtained from architectural descriptions/block diagrams, or estimated and varied for sensitivity analysis. By leveraging unique architectural properties of deep learning accelerators, we are able to systematically model a major class of hardware errors – transient errors in logic components – in software with high fidelity. Therefore, FIdelity is both quick and accurate, and does not require access to RTL.We thoroughly validate our FIdelity framework using Nvidia’s open-source accelerator called NVDLA, which shows that the results are highly accurate – out of 60K fault injection experiments, the software fault models derived using FIdelity closely match the behaviors observed from RTL simulations. Using the validated FIdelity framework, we perform a large-scale resilience study on NVDLA, which consists of 46M fault injection experiments running various representative deep neural network applications. We report the key findings and architectural insights, which can be used to guide the design of future accelerators.
Yi He 0010, Prasanna Balaprakash, Yanjing Li
MICRO3
2019 Protecting Page Tables from RowHammer Attacks using Monotonic Pointers in DRAM True-Cells
abstract
We identify an important asymmetry in physical DRAM cells that can be utilized to prevent RowHammer attacks by adding 18 lines of code to modify the OS memory allocator. Our small modification has a powerful impact on RowHammer's ability to bypass memory protection mechanisms and achieve a successful attack. Specifically, we identify two types of DRAM cells: true-cells and anti-cells. In a true-cell, a leaking capacitor will induce a '1'->'0' error, while in anti-cells, errors flow from '0'->'1'. We then create DRAM cell-type-aware memory allocation which enables a "monotonicity property" for a given data object. The monotonicity property is able to counter RowHammer attacks (and, to a broader extent, other memory attacks) by allocating only one type of cells for an object, thereby restricting error direction. We apply the monotonicity property to pointers in page tables by placing all page tables in true-cells that are above a "low water mark". We show that this approach successfully defends against page-table-based privilege escalation RowHammer attacks. Using established RowHammer-induced bit-flip error statistics, we provide proofs of the soundness and completeness of our technique and show that with our technique only one out of 2.04x10 5 systems is vulnerable to the attack, and the expected attack time on the vulnerable system is 231 days. We also provide application performance results from prototypes implemented through modifications to Linux kernels. Our cross-layer approach avoids undesirable energy cost, hardware changes, performance overhead, and high software complexity associated with prior countermeasures.
Xin-Chuan Wu, Timothy Sherwood, Fred Chong, Yanjing Li
ASPLOS4
2019 Cross-Layer Resilience: Challenges, Insights, and the Road Ahead
abstract
Resilience to errors in the underlying hardware is a key design objective for a large class of computing systems, from embedded systems all the way to the cloud. Sources of hardware errors include radiation, circuit aging, variability induced by manufacturing and operating conditions, manufacturing test escapes, and early-life failures. Many publications have suggested that cross-layer resilience, where multiple error resilience techniques from different layers of the system stack cooperate to achieve cost-effective resilience, is essential for designing cost-effective resilient digital systems. This paper presents a comprehensive overview of cross-layer resilience by addressing fundamental cross-layer resilience questions, by summarizing insights derived from recent advances in cross-layer resilience research, and by discussing future cross-layer resilience challenges.
Eric Cheng, Daniel Mueller-Gritschneder, Jacob A. Abraham, Pradip Bose, Alper Buyuktosunoglu, Deming Chen, Hyungmin Cho, Yanjing Li, Uzair Sharif, Kevin Skadron, Mircea R. Stan, Ulf Schlichtmann, Subhasish Mitra
DAC8
2019 Crash Skipping: A Minimal-Cost Framework for Efficient Error Recovery in Approximate Computing Environments
abstract
We present a lightweight technique to minimize error recovery costs in approximate computing environments. We take advantage of the key observation that if an application crashes in a "non-critical" region of its execution, then skipping the crash and allowing the execution to continue oftentimes results in "acceptable" output, due to the inherent fault-tolerance of approximate applications. By skipping application crashes, the program is given a chance to recover from an error on its own, without expending computing power towards error recovery. The system-level support required to implement our Crash Skipping technique imposes negligible overhead. Experimental results from representative approximate applications demonstrate that our technique is effective, resulting in successful error recovery for 56% of application crash cases on average, with a maximum recovery rate of 81%. By combining our technique with application restart, we obtain ~33% improvement in performance/energy consumption compared to recovering from crashes by restarting alone. This benefit is comparable to what can be achieved using aggressive checkpointing techniques, but without the significant costs in system design and complexity that such techniques impose.
Yan Verdeja Herms, Yanjing Li
ACM Great Lakes Symposium on VLSI2
2019 Two-Stream Multi-Task Network for Fashion Recognition
abstract
In this paper, we present a two-stream multi-task network for fashion recognition. This task is challenging as fashion clothing always contain multiple attributes, which need to be predicted simultaneously for real-time industrial systems. To handle these challenges, we formulate fashion recognition into a multi-task learning problem, including landmark detection, category and attribute classifications, and solve it with the proposed deep convolutional neural network. We design two knowledge sharing strategies which enable information transfer between tasks and improve the overall performance. The proposed model achieves state-of-the-art results on large-scale fashion dataset comparing to the existing methods, which demonstrates its great effectiveness and superiority for fashion recognition.
Peizhao Li, Yanjing Li, Xiantong Zhen
ICIP2
2019 Time-Slicing Soft Error Resilience in Microprocessors for Reliable and Energy-Efficient Execution
abstract
Resilience to soft errors is essential for ensuring the robustness of a computing system. In this paper, we present a new soft error resilience approach called TSSER (time-sliced soft error resilience), which enables resilience features for instructions that are most likely to cause errors only to minimize system-level energy costs while achieving high levels of resilience. Our TSSER idea (1) takes advantage of the observation that protecting a fraction of the instructions in an application already achieves most of the resilience benefits, (2) utilizes circuit-level features that allow resilience mode to be turned on/off, and (3) bridges the gap between application knowledge and circuit features by devising novel ISA and microarchitectural techniques to achieve optimized tradeoffs. Our results obtained from RTL implementation and detailed simulation show that, for various applications from the SPEC and PARSEC benchmark suites, TSSER achieves 65X reduction in SDC rate (a common metric to measure soft error resilience) while imposing 16.8% (11.3%) processor-level energy cost for in-order (out-of-order) processors. This is a significant improvement compared to existing techniques that impose 32% - 81% (18% -83 %) energy overhead. Our technique also enables flexible tradeoffs between SDC rate and system costs.
Yi He 0010, Yanjing Li
ITC2
2019 Direct-modulated optical networks for interposer systems
abstract
We present a new interposer-level optical network based on direct-modulated lasers such as vertical-cavity surface-emitting lasers (VCSELs) or transistor lasers (TLs). Our key observation is that, the physics of these lasers is such that they must transmit significantly more power (21x) than is needed by the receiver. We take advantage of this excess optical power to create a new network architecture called Rome, which splits optical signals using passive splitters to allow flexible bandwidth allocation among different transmitter and receiver pairs while imposing minimal power and design costs. Using multi-chip module GPUs (MCM-GPUs) as a case study, we thoroughly evaluate network power and performance, and show that (1) Rome is capable of efficiently scaling up MCM-GPUs with up to 1024 streaming multiprocessors, and (2) Rome outperforms various competing designs in terms of energy efficiency (by up to 4x) and performance (by up to 143%).
Mohammad Reza Jokar, Lunkai Zhang, John M. Dallesasse, Fred Chong, Yanjing Li
NOCS5
2017 Cross-layer refresh mitigation for efficient and reliable DRAM systems: A comparative study
abstract
DRAM is a crucial component in computing systems, and is expected to be even more important as data-intensive applications become more prominent. A key challenge in advancing DRAM technology is the growing cost of refresh operations, which can impose a large impact on the energy efficiency of DRAM modules. Existing refresh mitigation techniques all require hardware modifications, which may be undesirable. In this paper, we make two major contributions. First, we present a new cross-layer refresh mitigation approach that takes into account both DRAM retention time characteristics and application behaviors. The main ideas are: (1) uniformly lower refresh rate to improve DRAM energy efficiency, (2) utilize a combination of system-level memory tests and ECC/scrubbing to detect and correct errors resulting from the reduced refresh rate, and (3) perform software-based memory repair so that DRAM cells in which errors occur because of the reduced refresh rate are not mapped to application space. Our approach reduces refresh power by 98.6% and improves overall DRAM energy efficiency by up to 37.7% without sacrificing reliability, at the low price of a small reduction in main memory capacity. Second, we perform thorough experiments and analysis to compare our cross-layer approach with prior work. In general, our approach achieves better or similar power benefits, but it does not require any modifications to the hardware.
Xiaoan Ding, Xi Liang 0002, Yanjing Li
ITC3
2014 Special session 11C: Young professionals in test - Elevator talks
abstract
This session is organized as part of the activities sponsored by IEEE Test Technology Technical Council (TTTC) Young Professionals Forum. The primary goal of this forum is to align the young professionals, both from industry and academia, working in the broad domain of manufacturing test and applications, with the activities of TTTC. This forum was initiated in 2013 and had its first meetings at VLSI Test Symposium (VTS) and International Test Conference (ITC) of that year. This year we are expanding our presence in the VTS by introducing a new session showcasing the research conducted by the young colleagues from industry and academia. The Elevator Talk session includes presentations on “Malicious Aging Acceleration in Processors” by Naghmeh Karimi from New York Polytechnic University, “RF Built-In Test with Non-intrusive Sensors” by Haralampos Stratigopoulos from TIMA Laboratory, France, “Advanced Process Bring-up” by Sounil Biswas from nVidia, “Exploration of Vector-based Integer Arithmetic on Intel Xeon Processors” by Michail Maniatakos from New York University, Abu Dhabi and “Detecting Hardware Trojans with Self-Reference Timing Tests” by Eshan Singh from Intel Corporation. Through this Elevator Talk session, we encourage a broader section of young colleagues to participate in the similar session of future meetings.
Alodeep Sanyal, Yanjing Li
VTS2
2014 Special session 12C: Young professionals in test - Town meeting
abstract
In the year 2013, IEEE Test Technology Technical Council (TTTC) took an initiative to establish a forum involving young professionals (both from industry and academia) working in the broad domain of manufacturing test, diagnosis, debug, yield improvement and related areas. We organized panel meetings in conjunction with VLSI Test Symposium (VTS) and International Test Conference (ITC) last year to focus on the professional needs identified by the young colleagues that TTTC can help address. This is the third successive meeting of this forum that involves a diversified group of young professionals currently employed in the leading US semiconductor/EDA companies and in academia. The session will be held in town-hall format, organized by Dr. Alodeep Sanyal from Intel, and moderated by Dr. Yervant Zorian from Synopsys. The panelists will be involved in defining the charter for this newly-formed TTTC forum and establish a committee that will actively engage in monitoring these activities. Few of the focus areas include: (a) Create and maintain a webpage for Young Professionals (YP) Forum linked with the TTTC webpage, (b) Establish and nurture professional collaboration between young colleagues in academia and industry, (c) Create and maintain an employment opportunity database for graduating students. We invite everybody to attend the panel and voice their opinion from the audience.
Alodeep Sanyal, Yanjing Li, Yervant Zorian
VTS2
2014 Effective Post-Silicon Validation of System-on-Chips Using Quick Error Detection
abstract
This paper presents the Quick Error Detection (QED) technique for systematically creating families of post-silicon validation tests that quickly detect bugs inside processor cores and uncore components (cache controllers, memory controllers, and on-chip interconnection networks) of multicore system on chips (SoCs). Such quick detection is essential because long error detection latency, the time elapsed between the occurrence of an error due to a bug and its manifestation as an observable failure, severely limits the effectiveness of traditional post-silicon validation approaches. QED can be implemented completely in software, without any hardware modification. Hence, it is readily applicable to existing designs. Results using multiple hardware platforms, including the Intel® Core™ i7 SoC, and a state-of-the-art commercial multicore SoC, along with simulation results using an OpenSPARC T2-like multicore SoC with bug scenarios from commercial multicore SoCs demonstrate: 1) error detection latencies of post-silicon validation tests can be very long, up to billions of clock cycles, especially for bugs inside uncore components; 2) QED shortens error detection latencies by up to nine orders of magnitude to only a few hundred cycles for most bug scenarios; and 3) QED enables up to a fourfold increase in bug coverage.
Ted Hong, Yanjing Li, Eswaran S, Sharad Kumar, Farzan Fallah, Nagib Hakim, Donald S. Gardner, Subhasish Mitra
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2013 Overcoming post-silicon validation challenges through quick error detection (QED)
abstract
Existing post-silicon validation techniques are generally ad hoc, and their cost and complexity are rising faster than design cost. Hence, systematic approaches to post-silicon validation are essential. Our research indicates that many of the bottlenecks of existing post-silicon validation approaches are direct consequences of very long error detection latencies. Error detection latency is the time elapsed between the activation of a bug during post-silicon validation and its detection or manifestation as a system failure. In our earlier papers, we created the Quick Error Detection (QED) technique to overcome this significant challenge. QED systematically creates a wide variety of post-silicon validation tests to detect bugs in processor cores and uncore components of multi-core System-on-Chips (SoCs) very quickly, i.e., with very short error detection latencies. In this paper, we present an overview of QED and summarize key results: 1. Error detection latencies of “typical” post-silicon validation tests can range up to billions of clock cycles. 2. QED shortens error detection latencies by up to 6 orders of magnitude. 3. QED enables 2- to 4-fold improvement in bug coverage. QED does not require any hardware modification. Hence, it is readily applicable to existing designs.
Ted Hong, Yanjing Li, Farzan Fallah, Donald S. Gardner, Nagib Hakim, Subhasish Mitra
DATE3
2013 Self-repair of uncore components in robust system-on-chips: An OpenSPARC T2 case study
abstract
Self-repair replaces/bypasses faulty components in a system-on-chip (SoC) to keep the system functioning correctly even in the presence of permanent faults. Such faults may result from early-life failures, circuit aging, and manufacturing defects and variations. Unlike on-chip memories, processor cores, and networks-on-chip, little attention has been paid to self-repair of uncore components (e.g., cache controllers, memory controllers, and I/O controllers) that occupy significant portions of multi-core SoCs. In this paper, we present new techniques that utilize architectural features to achieve self-repair of uncore components while incurring low area, power, and performance costs. We demonstrate the effectiveness and practicality of our techniques, using the industrial OpenSPARC T2 SoC with 8 processor cores that support 64 hardware threads. Our key results are: 1. Our techniques enable effective self-repair of any single faulty uncore component with 7.5% post-layout chip-level area impact and 3% power impact. In contrast, existing redundancy techniques impose high (e.g., 16%) area costs. Our techniques do not incur any performance impact in fault-free systems. In the presence of a single faulty uncore component, there can be a 5% application performance impact. 2. Our techniques are capable of self-repairing multiple faulty uncore components without any additional area impact, but with graceful degradation of application performance. 3. Our techniques achieve high self-repair coverage of 97.5% in the presence of a single fault. Our self-repair techniques also enable flexible tradeoffs between self-repair coverage and area costs. For example, 75% self-repair coverage can be achieved with 3.2% post-layout chip-level area impact.
Yanjing Li, Eric Cheng, Samy Makar, Subhasish Mitra
ITC1
2010 Cross-layer error resilience for robust systems
abstract
A large class of robust electronic systems of the future must be designed to perform correctly despite hardware failures. In contrast, today's mainstream systems typically assume error-free hardware. Classical fault-tolerant computing techniques are too expensive for this purpose. This paper presents an overview of new techniques that can enable a sea change in the design of cost-effective robust systems. These techniques utilize globally-optimized cross-layer approaches, i.e., across device, circuit, architecture, runtime, and application layers, to overcome hardware failures.
Larkhoon Leem, Hyungmin Cho, Hsiao-Heng Lee, Young Moon Kim, Yanjing Li, Subhasish Mitra
ICCAD5
2010 QED: Quick Error Detection tests for effective post-silicon validation
abstract
Long error detection latency, the time elapsed between the occurrence of an error caused by a bug and its manifestation as a system-level failure, is a major challenge in post-silicon validation of robust systems. In this paper, we present a new technique called Quick Error Detection (QED), which transforms existing post-silicon validation tests into new validation tests that significantly reduce error detection latency. QED transformations allow flexible tradeoffs between error detection latency, coverage, and complexity, and can be implemented in software with little or no hardware changes. Results obtained from hardware experiments on quad-core Intel®Core™ i7 hardware platforms and from simulations on a multi-core MIPS processor design demonstrate that: 1. QED significantly improves error detection latencies by six orders of magnitude, i.e., from billions of cycles to a few thousand cycles or less. 2. QED transformations do not degrade the coverage of validation tests as estimated empirically by measuring the maximum operating frequencies over a wide range of operating voltage points. 3. QED tests improve coverage by detecting errors that escape the original non-QED tests.
Ted Hong, Yanjing Li, Sung-Boem Park, Diana Mui, Ziyad Abdel Kaleq, Nagib Hakim, Helia Naeimi, Donald S. Gardner, Subhasish Mitra
ITC2
2010 Concurrent autonomous self-test for uncore components in system-on-chips
abstract
Concurrent autonomous self-test, or online self-test, allows a system to test itself, concurrently during normal operation, with no system downtime visible to the end-user. Online self-test is important for overcoming major reliability challenges such as early-life failures and circuit aging in future System-on-Chips (SoCs). To ensure required levels of overall reliability of SoCs, it is essential to apply online self-test to uncore components, e.g., cache controllers, DRAM controllers, and I/O controllers, in addition to processor cores. This is because uncore components can account for a significant portion of the overall logic area of a multi-core SoC. In this paper, we present an efficient online self-test technique for uncore components in SoCs. We achieve extremely high test coverage by storing high-quality test patterns in off-chip non-volatile storage. However, a simple technique that stalls the uncore-component-under-test can result in significant system performance degradation or even visible system unresponsiveness. Our new techniques overcome these challenges and enable cost-effective online self-test of uncore components through three special hardware features: 1. resource reallocation and sharing (RRS); 2. no-performance-impact testing; and, 3. smart backups. Implementation of online self-test for uncore components of the open-source OpenSPARC T2 multi-core SoC, using a combination of these three techniques, achieves high test coverage at < 1% area impact, < 1% power impact, and < 3% system-level performance impact. These results demonstrate the effectiveness and practicality of our techniques.
Yanjing Li, Onur Mutlu, Donald S. Gardner, Subhasish Mitra
VTS1
2009 Operating system scheduling for efficient online self-test in robust systems
abstract
Very thorough online self-test is essential for overcoming major reliability challenges such as early-life failures and transistor aging in advanced technologies. This paper demonstrates the need for operating system (OS) support to efficiently orchestrate online self-test in future robust systems. Experimental data from an actual dual quad-core system demonstrate that, without software support, online self-test can significantly degrade performance of soft real-time and computation-intensive applications (by up to 190%), and can result in perceptible delays for interactive applications. To mitigate these problems, we develop OS scheduling techniques that are aware of online self-test, and schedule/migrate tasks in multi-core systems by taking into account the unavailability of one or more cores undergoing online self-test. These techniques eliminate any performance degradation and perceptible delays in soft real-time and interactive applications (otherwise introduced by online self-test), and significantly reduce the impact of online self-test on the performance of computation-intensive applications. Our techniques require minor modifications to existing OS schedulers, thereby enabling practical and efficient online self-test in real systems.
Yanjing Li, Onur Mutlu, Subhasish Mitra
ICCAD1
2008 CASP: Concurrent Autonomous Chip Self-Test Using Stored Test Patterns
abstract
CASP, concurrent autonomous chip self-test using stored test patterns, is a special kind of self-test where a system tests itself concurrently during normal operation without any downtime visible to the end-user. CASP consists of two ideas: 1. Storage of very thorough test patterns in non-volatile memory; and, 2. Architectural and system-level support for autonomous testing of one or more cores in a multi-core system using stored patterns, concurrently with normal system operation, without bringing down the entire system. CASP enables design of robust systems with built-in features for circuit failure prediction, error detection, self-diagnosis and self-repair. Such systems are necessary to overcome major reliability challenges in scaled-CMOS technologies. Implementation of CASP in the OpenSPARC Tl multi-core processor demonstrates its effectiveness and practicality.
Yanjing Li, Samy Makar, Subhasish Mitra
DATE1
2008 VAST: Virtualization-Assisted Concurrent Autonomous Self-Test
abstract
Virtualization-Assisted concurrent, autonomous Self-Test, or VAST, enables a multi-/many-core system to test itself, concurrently during normal operation, without any user-visible downtime. Such on-line self-test is required for large-scale robust systems with built-in support for circuit failure prediction, failure detection, diagnosis, and self-healing. The main idea behind VAST is hardware and software co-design of on-line self-test features in a multi-/many-core system through integration of: 1. multi-/many-core architecture, 2. virtualization software, and, 3. special self-test techniques such as BIST (Built-In Self-Test) or CASP (Concurrent Autonomous chip self-test using Stored Patterns). As a result, optimized trade-offs in system design complexity, system performance and power impact, and test thoroughness are possible. Experimental results from an actual multi-core system demonstrate that: 1. VAST is practical and effective; and, 2. Special VAST-supported self-test policies enable extremely thorough on-line self-test with very small performance impact.
Hiroaki Inoue, Yanjing Li, Subhasish Mitra
ITC2