Wei Li 0159

dblp:64/6025-159 · DBLP profile ↗
← Back
18ranked-venue papers
8as first author
12since 2021 · last 2025
0000-0001-9822-6021ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 8 first-author · 11 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Software engineering, systems software and programming languages · 2
YearPublicationVenuePosition
2025 IC-PEPR: PEPR Testing Goes Intra-Cell
abstract
Pseudo-Exhaustive Physically-Aware Region (PEPR) testing, in its most general application, rasterizes the layout of a logic circuit into overlapping three-dimensional regions of user-defined size. Faults that correspond to the exhaustive testing of each subcircuit within a region are defined and used for automatic test pattern generation (ATPG), fault simulation, and diagnosis. Evaluation of tens of thousands of chip failures demonstrated the effectiveness of PEPR in capturing the exact behavior of defects.However, deployment of PEPR is challenged by the large number of faults produced for exhaustively testing each region. Analysis revealed that there are a small percentage of regions that require a significant number of faults. For instance, a 14nm test chip contains regions that require more than 32M faults to test. While it is certainly possible to define a region size that produces large subcircuits, we found that large subcircuits predominantly result from including the entirety of a cell when a region simply intersects a small portion of the cell. To remedy this situation, we have developed IC-PEPR, a novel intra-cell extension to PEPR that significantly reduces the number of resulting faults. Specifically, by exploiting equivalence within a cell, intra-cell components within a region are controlled to all possible values without applying all cell-level input patterns. Applying IC-PEPR testing reduces the number of faults for a commercial benchmark circuit by more than 100X. This reduction in fault count results in a corresponding reduction in ATPG run time (almost 50X reduction) and test set size (>10X reduction).
Chris Nigh, Ruben Purdy, Wei Li 0159, Subhasish Mitra, R. D. (Shawn) Blanton
ITC3
2025 VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning
abstract
Recent advancements in vision-language models (VLMs) have improved performance by increasing the number of visual tokens, which are often significantly longer than text tokens. However, we observe that most real-world scenarios do not require such an extensive number of visual tokens. While the performance drops significantly in a small subset of OCR-related tasks, models still perform accurately in most other general VQA tasks with only 1/4 resolution. Therefore, we propose to dynamically process distinct samples with different resolutions, and present a new paradigm for visual token compression, namely, VisionThink. It starts with a downsampled image and smartly decides whether it is sufficient for problem solving. Otherwise, the model could output a special token to request the higher-resolution image. Compared to existing Efficient VLM methods that compress tokens using fixed pruning ratios or thresholds, VisionThink autonomously decides whether to compress tokens case by case. As a result, it demonstrates strong fine-grained visual understanding capability on OCR-related tasks, and meanwhile saves substantial visual tokens on simpler tasks. We adopt reinforcement learning and propose the LLM-as-Judge strategy to successfully apply RL to general VQA tasks. Moreoever, we carefully design a reward function and penalty mechanism to achieve a stable and reasonable image resize call ratio. Extensive experiments demonstrate the superiority, efficiency, and effectiveness of our method. All our code and data are open-sourced.
Senqiao Yang, Wei Li 0159, Zejun Ma 0001, Bei Yu 0001, Hengshuang Zhao, Jiaya Jia
NeurIPS5
2025 CHEF: CHaracterizing Elusive Logic Circuit Failures
abstract
Logic circuit diagnosis is an essential tool for improving manufacturing yield. However, there is a significant disparity between the behavior predicted by conventional fault models and the actual behavior observed in failing circuits. This undermines the performance of diagnosis methodologies that rely on conventional fault models to characterize defect behavior. Recently, a parameterizable test metric called PEPR has demonstrated the ability to precisely bound defect behavior, even when it deviates from conventional fault model assumptions. This paper describes a novel diagnosis methodology that uses the PEPR metric to determine (i) the physical location of a defect and (ii) the precise changes to the logic functionality of the affected circuit in the form of a custom fault model. The methodology, called CHEF, is applied to over 700 fail logs from a 22nm industrial test chip. Results demonstrate CHEF precisely characterizes defects that conventional diagnosis cannot, specifically when failure behavior does not align with a conventional fault model. Furthermore, CHEF achieves a significant increase in diagnostic resolution compared to conventional diagnosis, more than doubling the number of failures with a physical resolution better than 1µm2. Finally, diagnostic patterns generated by CHEF demonstrate the capability to further refine defect characterization.
Ruben Purdy, Chris Nigh, Wei Li 0159, R. D. (Shawn) Blanton
VTS3
2024 DGR: Differentiable Global Router
abstract
Modern VLSI design flows necessitate fast and high-quality global routers. In this paper, we introduce DGR, a differentiable global router capable of concurrent optimization for hundreds of thousands of nets 1. Our innovation lies in the development of a routing Directed Acyclic Graph (DAG) forest to represent the 2D pattern routing space for all nets, enabling coordinated selection of Steiner trees and 2-pin routing paths from a global perspective. For efficient search within the DAG forest, we relax the discrete search space to be continuous and develop a differentiable solver accelerated by deep learning toolkits on GPUs. Experimental results demonstrate that DGR substantially mitigates routing overflow while concurrently reducing total wirelengths from 0.95% to 4.08% and via numbers from 1.28% to 2.54% in congested testcases compared to state-of-the-art academic global routers. Additionally, DGR exhibits favorable scalability in both runtime and memory with respect to the number of nets.
Wei Li 0159, Rongjian Liang, Anthony Agnesina, Chia-Tung Ho, Anand Rajaram, Haoxing Ren
DAC1
2024 Silent Data Corruption: Test or Reliability Problem?
abstract
Recently, companies such as Google, Meta (Facebook), and Microsoft reported in the mainstream press about seemingly random errors which, initially undetected ("silently"), had crept into their large cloud data centers. These reports mentioned that very specific instructions were intermittently incorrectly executed, propagated through the operating system, and would potentially manifest themselves as application-level errors. Are the root causes of these so-called silent data errors test escapes and/or reliability issues? Why are they only noticed now? Is that only the case because such large server farms bring together larger numbers of CPUs than ever seen before? And what counter measures can we take against them?
Erik Jan Marinissen, Harish Dattatraya Dixit, R. D. (Shawn) Blanton, Aaron Kuo, Wei Li 0159, Subhasish Mitra, Chris Nigh, Ruben Purdy, Ben Kaczer, Dishant Sangani, Pieter Weckx, Philippe Roussel, Georges Gielen
ETS5
2024 Faulty Function Extraction for Defective Circuits
abstract
It is well-known that understanding the behavior of silicon failures is an essential step in yield learning. It is also becoming more important for producing high-quality silicon due to the increasing number of defects detected fortuitously. In order to meet this need, a new approach for extracting the precise faulty function from defective logic circuits is described. The approach is applied to nearly a 1,000 14nm failures and one use case of the results on improving ATPG is discussed.
Chris Nigh, Ruben Purdy, Wei Li 0159, Subhasish Mitra, R. D. (Shawn) Blanton
ETS3
2023 Global Floorplanning via Semidefinite Programming
abstract
A major task in chip design involves identifying the location and shape of each major design block/module in the layout footprint. This is commonly known as floorplanning. The first step of this task is known as global floorplanning and involves identifying a location for each module that minimizes wire length and leaves sufficient area for each module. Existing global floorplanning methods either have a non-convex problem formulation, or have trivial global solutions with no guarantee on the quality of the result. Here, we model the global floorplanning as a Semi-Definite Programming (SDP) problem with a rank constraint. We replace the rank constraint with a direction matrix and convexify the problem, whose solution is shown to be a global optimum if an appropriate direction matrix is chosen. To calculate the direction matrix, a convex iteration algorithm is used where the problem is decomposed into two SDP sub-problems. Furthermore, we introduce a series of techniques that enhance the flexibility, accuracy, and efficiency of our algorithm. Design experiments demonstrate that our proposed method reduces the average wirelength up to 20% for different benchmarks and outline aspect ratios.
Wei Li 0159, José M. F. Moura, R. D. (Shawn) Blanton
DAC1
2022 PEPR: Pseudo-Exhaustive Physically-Aware Region Testing
abstract
Recent reports indicate that existing fault models and test metrics result in substantial manufacturing test escapes that cause major system-level challenges such as silent data corruption resulting from incorrect computations. Such test escapes are often detected today after system deployment (e.g., in the field) using a variety of synthetic and application workloads. In this work, a new test metric is investigated for detecting defects that escape existing test approaches. PEPR (Pseudo-Exhaustive Physically-Aware Region) testing comprehensively analyzes both the physical layout and the logic netlist to identify single- or multi- output sub-circuits. The resulting sub-circuits are exhaustively tested to detect timing-independent combinational (TIC) defects. Analyses demonstrate that PEPR-based scan tests detect TIC defects perfectly (100%) when examining fail data from over 30,000 14nm failing chips. In contrast, existing fault models and test metrics might result in up to 95 % of TIC defects being detected fortuitously. Strategies for addressing increased test pattern count resulting from the pseudo-exhaustive nature of PEPR testing are also discussed.
Wei Li 0159, Chris Nigh, Danielle Duvalsaint, Subhasish Mitra, R. D. (Shawn) Blanton
ITC1
2022 Adaptive Layout Decomposition With Graph Embedding Neural Networks
abstract
Multiple patterning layout decomposition (MPLD) has been widely investigated, but so far there is no decomposer that dominates others in terms of both result quality and efficiency. This observation motivates us to explore how to adaptively select the most suitable MPLD strategy for a given layout graph, which is nontrivial and still an open problem. In this article, we propose a layout decomposition framework based on graph convolutional networks to obtain the graph embeddings of the layout. The graph embeddings are used for graph library construction, decomposer selection, graph matching, stitch removal prediction, and graph coloring. In addition, we design a fast nonstitch layout decomposition algorithm that purely depends on the message passing graph neural network. The experimental results show that our graph embedding-based framework can achieve optimal decompositions in the widely used benchmark with a significant runtime drop even compared with fast but nonoptimal heuristics.
Wei Li 0159, Yuzhe Ma, Yibo Lin, Bei Yu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2021 TreeNet: Deep Point Cloud Embedding for Routing Tree Construction
abstract
In the routing tree construction, both wirelength (WL) and path-length (PL) are of importance. Among all methods, PD-II and SALT are the two most prominent ones. However, neither PD-II nor SALT always dominates the other one in terms of both WL and PL for all nets. In addition, estimating the best parameters for both algorithms is still an open problem. In this paper, we model the pins of a net as point cloud and formalize a set of special properties of such point cloud. Considering these properties, we propose a novel deep neural net architecture, TreeNet, to obtain the embedding of the point cloud. Based on the obtained cloud embedding, an adaptive workflow is designed for the routing tree construction. Experimental results show that the proposed TreeNet is superior to other mainstream models for the point cloud on classification tasks. Moreover, the proposed adaptive workflow for the routing tree construction outperforms SALT and PD-II in terms of both efficiency and effectiveness.
Wei Li 0159, Yuxiao Qu, Gengjie Chen, Yuzhe Ma, Bei Yu 0001
ASP-DAC1
2021 Learning Point Clouds in EDA
abstract
The exploding of deep learning techniques have motivated the development in various fields, including intelligent EDA algorithms from physical implementation to design for manufacturability. Point cloud, defined as the set of data points in space, is one of the most important data representations in deep learning since it directly pre- serves the original geometric information without any discretization. However, there are still some challenges that stifle the applications of point clouds in the EDA field. In this paper, we first review previous works about deep learning in EDA and point clouds in other fields. Then, we discuss some challenges of point clouds in EDA raised by some intrinsic characteristics of point clouds. Finally, to stimulate future research, we present several possible applications of point clouds in EDA and demonstrate the feasibility by two case studies.
Wei Li 0159, Guojin Chen, Ran Chen 0001, Bei Yu 0001
ISPD1
2021 OpenMPL: An Open-Source Layout Decomposer
abstract
Multiple patterning lithography has been widely adopted in advanced technology nodes of VLSI manufacturing. As a key step in the design flow, multiple patterning layout decomposition (MPLD) is critical to design closure. Due to the$\mathcal {N} \mathcal {P} $-hardness of the general decomposition problem, various efficient algorithms have been proposed with high-quality solutions. However, with increasingly complicated design flow and peripheral processing steps, developing a high-quality layout decomposer becomes more and more difficult, slowing down further advancement in this field. This article presents$\mathsf {OpenMPL}$(2020), an open-source layout decomposition framework, with well-separated peripheral processing and core solving steps. Besides, previous algorithms or techniques are inspected and several issues are discovered. We then propose corresponding new algorithms to resolve these issues. The experiments demonstrate the effectiveness of our proposed algorithms and the efficiency of$\mathsf {OpenMPL}$.
Wei Li 0159, Yuzhe Ma, Qi Sun 0002, Yibo Lin, Iris Hui-Ru Jiang, Bei Yu 0001, David Z. Pan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2020 Adaptive Layout Decomposition with Graph Embedding Neural Networks
abstract
Multiple patterning lithography decomposition (MPLD) has been widely investigated, but so far there is no decomposer that dominates others in terms of both the optimality and the efficiency. This observation motivates us exploring how to adaptively select the most suitable MPLD strategy for a given layout graph, which is non-trivial and still an open problem. In this paper, we propose a layout decomposition framework based on graph convolutional networks to obtain the graph embeddings of the layout. The graph embeddings are used for graph library construction, decomposer selection and graph matching. Experimental results show that our graph embedding based framework can achieve optimal decompositions under negligible runtime overhead even comparing with fast but non-optimal heuristics.
Wei Li 0159, Jialu Xia, Yuzhe Ma, Yibo Lin, Bei Yu 0001
DAC1
2020 DeepBillboard: systematic physical-world testing of autonomous driving systems
abstract
Deep Neural Networks (DNNs) have been widely applied in autonomous systems such as self-driving vehicles. Recently, DNN testing has been intensively studied to automatically generate adversarial examples, which inject small-magnitude perturbations into inputs to test DNNs under extreme situations. While existing testing techniques prove to be effective, particularly for autonomous driving, they mostly focus on generating digital adversarial perturbations, e.g., changing image pixels, which may never happen in the physical world. Thus, there is a critical missing piece in the literature on autonomous driving testing: understanding and exploiting both digital and physical adversarial perturbation generation for impacting steering decisions. In this paper, we propose a systematic physical-world testing approach, namely DeepBillboard, targeting at a quite common and practical driving scenario: drive-by billboards. DeepBillboard is capable of generating a robust and resilient printable adversarial billboard test, which works under dynamic changing driving conditions including viewing angle, distance, and lighting. The objective is to maximize the possibility, degree, and duration of the steering-angle errors of an autonomous vehicle driving by our generated adversarial billboard. We have extensively evaluated the efficacy and robustness of DeepBillboard by conducting both experiments with digital perturbations and physical-world case studies. The digital experimental results show that DeepBillboard is effective for various steering models and scenes. Furthermore, the physical case studies demonstrate that DeepBillboard is sufficiently robust and resilient for generating physical-world adversarial billboard tests for real-world driving under various weather conditions, being able to mislead the average steering angle error up to 26.44 degrees. To the best of our knowledge, this is the first study demonstrating the possibility of generating realistic and continuous physical-world tests for practical autonomous driving systems; moreover, DeepBillboard can be directly generalized to a variety of other physical entities/surfaces along the curbside, e.g., a graffiti painted on a wall.
Husheng Zhou, Wei Li 0159, Zelun Kong, Yuqun Zhang, Bei Yu 0001, Lingming Zhang 0001, Cong Liu 0005
ICSE2
2020 Understanding Graphs in EDA: From Shallow to Deep Learning
abstract
As the scale of integrated circuits keeps increasing, it is witnessed that there is a surge in the research of electronic design automation (EDA) to make the technology node scaling happen. Graph is of great significance in the technology evolution since it is one of the most natural ways of abstraction to many fundamental objects in EDA problems like netlist and layout, and hence many EDA problems are essentially graph problems. Traditional approaches for solving these problems are mostly based on analytical solutions or heuristic algorithms, which require substantial efforts in designing and tuning. With the emergence of the learning techniques, dealing with graph problems with machine learning or deep learning has become a potential way to further improve the quality of solutions. In this paper, we discuss a set of key techniques for conducting machine learning on graphs. Particularly, a few challenges in applying graph learning to EDA applications are highlighted. Furthermore, two case studies are presented to demonstrate the potential of graph learning on EDA applications.
Yuzhe Ma, Zhuolun He, Wei Li 0159, Bei Yu 0001
ISPD3
2019 FIT: Fill Insertion Considering Timing
abstract
Dummy fill insertion is a mandatory step in modern semiconductor manufacturing process to reduce dielectric thickness variation, and provide nearly uniform pattern density for the chemical mechanical planarization (CMP) process. However, with the continuous shrinking of the VLSI technology nodes, the coupling effects between the inserted metal fills and signal tracks can severely affect the original timing closure of the layout design. In this paper, we propose a robust, efficient and high-performance framework for timing-aware dummy fill insertion, which simultaneously minimizes the coupling capacitance of critical signal wires and other wires. The experimental results on IC/CAD 2018 contest benchmarks shows that our proposed framework outperforms contest winner by 8% on critical coupling capacitance with 3.3× runtime speedup.
Bentian Jiang, Xiaopeng Zhang 0009, Ran Chen 0001, Gengjie Chen, Peishan Tu, Wei Li 0159, Evangeline F. Y. Young, Bei Yu 0001
DAC6
2019 A Unified Approximation Framework for Compressing and Accelerating Deep Neural Networks
abstract
Deep neural networks (DNNs) have achieved significant success in a variety of real world applications, i.e., image classification. However, tons of parameters in the networks restrict the efficiency of neural networks due to the large model size and the intensive computation. To address this issue, various approximation techniques have been investigated, which seek for a light weighted network with little performance degradation in exchange of smaller model size or faster inference. Both low-rankness and sparsity are appealing properties for the network approximation. In this paper we propose a unified framework to compress the convolutional neural networks (CNNs) by combining these two properties, while taking the nonlinear activation into consideration. Each layer in the network is approximated by the sum of a structured sparse component and a low-rank component, which is formulated as an optimization problem. Then, an extended version of alternating direction method of multipliers (ADMM) with guaranteed convergence is presented to solve the relaxed optimization problem. Experiments are carried out on VGG-16, AlexNet and GoogLeNet with large image classification datasets. The results outperform previous work in terms of accuracy degradation, compression rate and speedup ratio. The proposed method is able to remarkably compress the model (with up to 4.9X reduction of parameters) at a cost of little loss or without loss on accuracy.
Yuzhe Ma, Ran Chen 0001, Wei Li 0159, Fanhua Shang, Wenjian Yu, Minsik Cho, Bei Yu 0001
ICTAI3
2019 DeepFL: integrating multiple fault diagnosis dimensions for deep fault localization
abstract
Learning-based fault localization has been intensively studied recently. Prior studies have shown that traditional Learning-to-Rank techniques can help precisely diagnose fault locations using various dimensions of fault-diagnosis features, such as suspiciousness values computed by various off-the-shelf fault localization techniques. However, with the increasing dimensions of features considered by advanced fault localization techniques, it can be quite challenging for the traditional Learning-to-Rank algorithms to automatically identify effective existing/latent features. In this work, we propose DeepFL, a deep learning approach to automatically learn the most effective existing/latent features for precise fault localization. Although the approach is general, in this work, we collect various suspiciousness-value-based, fault-proneness-based and textual-similarity-based features from the fault localization, defect prediction and information retrieval areas, respectively. DeepFL has been studied on 395 real bugs from the widely used Defects4J benchmark. The experimental results show DeepFL can significantly outperform state-of-the-art TraPT/FLUCCS (e.g., localizing 50+ more faults within Top-1). We also investigate the impacts of deep model configurations (e.g., loss functions and epoch settings) and features. Furthermore, DeepFL is also surprisingly effective for cross-project prediction.
Xia Li 0009, Wei Li 0159, Yuqun Zhang, Lingming Zhang 0001
ISSTA2