Qi Xu 0004

dblp:84/1680-4 · DBLP profile ↗
← Back
44ranked-venue papers
11as first author
32since 2021 · last 2026
0000-0002-0375-9800ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 41 · 10 first-author · 29 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Real-Time Compensation Framework for Large-Scale ReRAM-Based Sparse LU Factorization
abstract
Recently, resistive switching random access memory (ReRAM)-based hardware accelerators have demonstrated unprecedented performance compared to digital accelerators. However, due to limitations in the manufacturing process and largescale integration, several significant non-ideal effects, including IR-Drop, Stuck-At-Fault, and device noises in real ReRAM-based crossbar arrays, are typically incurred. These non-ideal effects degrade signal integrity and performance, particularly in crossbar structures used for building high-density ReRAMs. Therefore, finding a fast and efficient software solution that can predict the effects of IR-drop without involving expensive hardware is highly desirable. In this work, addressing the main limitations of existing simulation methods, such as slow speed and high resource costs, we propose an efficient analysis of large-scale ReRAM crossbar arrays and the corresponding non-ideal factors based on sparse matrix modeling. We classify non-ideal factors into linear (e.g., IR-drop) and nonlinear categories (e.g., shot noise). For linear factors, super-nodal sparse LU factorizations are used to solve. The array-level results show that compared to SPICE simulation, our method achieves a numerical solution accuracy of 10.15 with 506.8 1253.3× faster and 17.46 42934.3× reduced memory usage. For nonlinear factors, we propose two solutions based on different requirements. In one method, we obtain an approximate initial solution by solving a linear system while disregarding the nonlinear contributions and subsequently apply an extended Anderson acceleration method to solve the nonlinear equation, which is suitable for high-precision solutions. Another method simplifies the nonlinear equation into an equivalent linear form. Theoretical validation confirms the effectiveness of this method, significantly enhancing simulation speed while maintaining accuracy. Moreover, we build a high-precision ReRAM accelerator architecture with real-time compensation. Experimental results demonstrate that the proposed architecture effectively mitigates accuracy loss caused by non-ideal factors.
Zaitian Chen, Bei Yu 0001, Song Chen 0001, Yi Kang, Qi Xu 0004
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2026 Efficient Routing-Based Synthesis for Digital Microfluidic Biochips via Reinforcement Learning
abstract
The use of digital microfluidic biochips (DMFBs) has highlighted their superiority in automatically executing biochemical assays by controlling tiny nano/picoliter droplets, which are moved in parallel to enhance throughput. Routing-based synthesis for DMFBs yields faster assay execution times compared to module-based synthesis when on-chip resource constraints are stringent. However, without predefining modules, it is very challenging to handle all the droplets directly on the chip for successfully executing the desired biochemical assay, especially in dynamic environments. Through modeling routing-based synthesis into two kinds of real-time decision tasks, i.e., transportation and mixing, this paper proposes a new routing-based synthesis framework that uses deep reinforcement learning (DRL) to train transportation and mixing agents respectively. Additionally, we design effective partial observations and curriculum learning (CL) schemes for both kinds of agents to improve their generalization ability and accelerate the training process. Compared to the state-of-the-art heuristic routing-based synthesis methods, more efficient synthesis processes of the given assays can be achieved using the proposed method of combining DRL and CL. For example, the average completion time on several real-world bioassay benchmarks (PCR, INVITRO, and PROTEIN) was reduced by 12.9% 18.5% approximately.
Qi Xu 0004, Hailong Yao 0002, Tsung-Yi Ho, Bo Yuan 0006
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2026 Attention-Based EDA Tool Parameter Explorer: From Hybrid Parameters to Multi-QoR Metrics
Donger Luo, Qi Sun 0002, Peng Xu 0052, Su Zheng, Qi Xu 0004, Tinghuan Chen, Bei Yu 0001, Hao Geng
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2026 Scalable High-Fidelity Solver for Large-Scale ReRAM Crossbar Arrays Under I-V Nonlinearity
abstract
Large-scale resistive random access memory (ReRAM) crossbar arrays have attracted considerable interest for in-memory computing (IMC) applications due to their high integration density and intrinsic parallelism. To enable systematic exploration of architectural design spaces, accurate modeling and efficient simulation of arrays are crucial. However, as arrays sizes increase, non-ideal effects—such as IR-Drop, I-V nonlinearity, and device noises—significantly degrade computational accuracy and efficiency. Although SPICE-based circuit simulators provide high fidelity, their excessive computational and memory overhead makes them impractical for simulating large-scale arrays under non-ideal conditions. In this article, we propose an efficient and scalable numerical framework for simulating large-scale ReRAM crossbar arrays under various conditions, including ideal behavior, I-V nonlinearity, and device noises, and so on. The proposed methodology integrates Cholesky decomposition with a fast iterative solver to enhance computational efficiency. Experimental results demonstrate that compared with existing solvers, our framework achieves high accuracy while greatly reducing runtime and memory consumption in modeling large-scale ReRAM crossbar arrays. This advantage is particularly evident under ReRAM nonlinearity and IR-Drop effects, achieving an average speedup of 162.9× over HSPICE across array sizes ranging from 128 to 2048. This work facilitates efficient and accurate design space exploration for next-generation ReRAM-based accelerators.
Qi Xu 0004, Song Chen 0001, Yi Kang, Bei Yu 0001
ACM Trans. Design Autom. Electr. Syst.2
2026 GoSteiner: Constructing Rectilinear Steiner Minimum Tree on Directed Graph
abstract
The Rectilinear Steiner Minimum Tree (RSMT) problem is a key issue in the back-end physical design of integrated circuits (ICs), which directly affects the quality of the routing. In this work, we formulate the RSMT problem as a sequential decision problem to develop an Actor-Critic reinforcement learning framework named GoSteiner. We utilize a directed graph representation method called GST for RSMT. Additionally, we introduce the Delaunay triangulation graph (DT) and sequentially construct GST on DT to solve the RSMT problem. An edge-aware graph attention network (EGAT) is designed to effectively encode the DT graph and the GST, while a transformer-based decoder is built to output policy. Furthermore, we propose a heuristic method to break high-degree nets (> 50 degrees). This approach fully leverages the critic’s ability to accurately estimate the wirelength of nets, significantly enhancing the quality and efficiency in high-degree nets construction. Experimental results demonstrate that compared with the exact algorithm GeoSteiner, GoSteiner only introduces ≤ 0.24% wirelength error on the ISPD18/19 benchmarks with million of nets. Meanwhile, the runtime for net within 500 pins is less than 62.37 ms .
Qi Xu 0004, Song Chen 0001, Yi Kang, Bei Yu 0001
ACM Trans. Design Autom. Electr. Syst.2
2025 ThePlace: Thermal-Aware Placement With Operator Learning-Based Ultra-Fast Simulator
abstract
Thermal issues are major concerns in integrated circuits (ICs) design. Typically, high temperature induces stress and carrier mobility changes between different materials, causing timing and reliability challenges in chip. In this paper, we propose a thermal-aware placement engine named ThePlace. It consists of an ultra-fast thermal simulation model using Fourier neural operator (FNO) to solve the steady-state heat conduction equation, followed by a force-directed global placement algorithm to co-optimize the peak temperature and wirelength in placements. The experimental results indicate that compared with the wirelength-driven placement approach DREAMPlace, ThePlace method enables significant temperature reduction with subtle variation in wirelength.
Xinfei Liu, Siting Liu 0002, Bei Yu 0001, Song Chen 0001, Qi Xu 0004
ASP-DAC5
2025 AIPlace: Analog IC Placement with Multi-Task Learning Framework
abstract
Layout design of analog integrated circuits is a time-consuming manual process with limited automation methods. Recently, advances in machine learning have opened up possibilities for automated design, making it a viable option to improve efficiency. In this paper, we present an innovative and highly effective approach to achieve automated analog circuit placement. We transform the analog placement constraints into multiple task objectives, and apply multi-task neural network learning to perform accurate placement solutions efficiently. Besides, the global position information is utilized to achieve more orderly placement. Due to the computational properties of the network, the method exhibits versatility in accommodating diverse scales of circuit netlists. Moreover, the model is trained through unsupervised learning. Compared to the supervised counterpart using many generated synthetic layout datasets, the proposed approach dramatically reduces the cost of placement data. Experimental results demonstrate that compared to SOTA works, the proposed placement learning method can achieve significant performance gains.
Jing Wang 0131, Song Chen 0001, Qi Xu 0004
ASP-DAC4
2025 LMM-IR: Large-Scale Netlist-Aware Multimodal Framework for Static IR-Drop Prediction
abstract
Static IR drop analysis is a fundamental and critical task in the field of chip design. Nevertheless, this process can be quite time-consuming, potentially requiring several hours. Moreover, addressing IR drop violations frequently demands iterative analysis, thereby causing the computational burden. Therefore, fast and accurate IR drop prediction is vital for reducing the overall time invested in chip design. In this paper, we firstly propose a novel multimodal approach that efficiently processes SPICE files through large-scale netlist transformer (LNT). Our key innovation is representing and processing netlist topology as 3D point cloud representations, enabling efficient handling of netlist with up to hundreds of thousands to millions nodes. All types of data, including netlist files and image data, are encoded into latent space as features and fed into the model for static voltage drop prediction. This enables the integration of data from multiple modalities for complementary predictions. Experimental results demonstrate that our proposed algorithm can achieve the best F1 score and the lowest MAE among the winning teams of the ICCAD 2023 contest and the state-of-theart algorithms.
Zhen Wang 0030, Hongquan He, Qi Xu 0004, Tinghuan Chen, Hao Geng
DAC4
2025 Robust and Efficient Adversarial Defense in SNNs via Image Purification and Joint Detection
abstract
Spiking neural networks (SNNs) leverage neural spikes to provide solutions for low-power intelligent applications on neuromorphic hardware. Although the spiking mechanism significantly enhances computational efficiency, especially in energy-constrained environments, they still lack resistance to noise perturbations and adversarial attacks. In this paper, we propose a defense framework based entirely on SNNs and design a fast and efficient training algorithm. The framework is divided into an image purification module and an adversarial detection module. The image purification module is employed for the extraction of noise and the reconstruction of input images. The adversarial detection module is utilized to differentiate between clean and adversarial images, thereby further enhancing defense performance. Meanwhile, our approach is highly flexible and can be seamlessly integrated with other defense strategies. Experimental results demonstrate that the proposed methodology outperforms state-of-the-art baselines in terms of defense effectiveness, training time and resource consumption.
Qi Xu 0004
ICASSP2
2024 OTPlace-Vias: A Novel Optimal Transport Based Method for High Density Vias Placement in 3D Circuits
abstract
Three-dimensional integrated circuit (3D IC) is an important manufacturing technology. In particular, the Monolithic 3D (M3D) technology stands out as a cutting-edge approach that provides higher integration density. However, M3D also introduces several challenges in terms of high density and computational complexity. In this paper, we propose a new approach for solving the inter-tier vias placement problem through optimal transport, which can be efficiently implemented in parallel with GPUs and consequently achieves significant speedup. Moreover, comparing with previous methods, our approach can also facilitate the processing of high integration density circuits to be more effective.
Qi Xu 0004, Hu Ding 0003
DAC2
2024 Parallel Multi-Objective Bayesian Optimization Framework for CGRA Microarchitecture
abstract
Recently, due to the flexibility and reconfigurability of Coarse-Grained Reconfigurable Architecture (CGRA), CGRA microarchitecture has become an inevitable trend to accelerate the convolution calculation in diverse deep neural networks. However, since the vast microarchitecture design space and the complicated VLSI verification flow, it is a huge challenge to explore a perfect microarchitecture to compromise between multiple performance metrics. In this paper, we formulate the CGRA microarchitecture design as a design space exploration problem, and propose a parallel multi-objective Bayesian optimization framework (PAMBOF) to automatically explore the CGRA microarchitecture design space. Meanwhile, high-precision performance and area models are built to enable fast design space exploration. To approximate the black-box objective function in the design space, the PAMBOF framework first builds multiple Gaussian processes (GP) with deep regularization kernel learning functions (DRKL-GP). Then a parallel Bayesian optimization algorithm is developed to sample a batch of candidate design points, which are simulated in parallel by the performance and area models. Experimental results demonstrate that compared to the prior arts, the proposed PAMBOF framework can search for a CGRA microarchitecture design with the better area and performance in a shorter runtime.
Wendi Sun, Xiaobing Ni, Kaixuan He, Qi Xu 0004, Song Chen 0001, Yi Kang
DATE5
2024 Attention-Based EDA Tool Parameter Explorer: From Hybrid Parameters to Multi-QoR metrics
abstract
Improving the outcomes of very-large-scale integration design without altering the underlying design enablement, such as process, device, interconnect, and IPs, is critical for integrated circuit (IC) designers. Parameter tuning for electronic design automation (EDA) tools is an emerging technology for improving the final design Quality-of-Result (QoR). However, many complex heuristics have been accreted upon previous complex heuristics integrated into tools, resulting in a vast number of tunable parameters. Even worse, these parameters include both continuous and discrete ones, making the parameter tuning process laborious and challenging. In this paper, we propose an attention-based EDA tool parameter explorer. A self-attention mechanism is developed to navigate the parameter importance. A hybrid space Gaussian process model is leveraged to optimize continuous and discrete parameters jointly, capturing their complex interactions. In addition, considering multiple QoR metrics and the large amount of time required to invoke EDA tools, a customized acquisition function based on expected hypervolume improvement (EHVI) is proposed to enable multi-objective optimization and parallel evaluation. Experimental results on a set of IWLS2005 benchmarks demonstrate the effectiveness and efficiency of our method.
Donger Luo, Qi Sun 0002, Qi Xu 0004, Tinghuan Chen, Hao Geng
DATE3
2024 Miracle: Multi-Action Reinforcement Learning-Based Chip Floorplanning Reasoner
abstract
Floorplanning is one of the most critical but time-consuming tasks in the chip design process. Machine learning techniques, especially reinforcement learning, have provided a promising direction for floorplanning design. In this paper, an end-to-end reinforcement learning (RL) framework is proposed to learn a policy for floorplanning automatically, in the combination of edge-augmented graph attention network (EGAT), position-wise multi-layer perceptron, and gated self-attention mechanism. We formulate floorplanning as a Markov Decision Process (MDP) model, where a multi-action mechanism and a dense reward function are developed to adapt the floorplanning problem. In addition, in order to make full use of prior knowledge, we further propose a supervised learning approach on the generated synthetic netlist-floorplan dataset. Experimental results demonstrate that, compared with state-of-the-art floorplanners, the proposed end-to-end framework significantly reduces wirelength with a smaller area.
Qi Xu 0004, Hao Geng, Song Chen 0001, Yi Kang
DATE2
2024 NicePIM: Design Space Exploration for Processing-In-Memory DNN Accelerators With 3-D Stacked-DRAM
abstract
With the widespread use of deep neural networks (DNNs) in intelligent systems, DNN accelerators with high performance and energy efficiency are greatly demanded. As one of the feasible processing-in-memory (PIM) architectures, 3D-stacked-DRAM-based PIM (DRAM-PIM) architecture enables large-capacity memory and low-cost memory access, which is a promising solution for DNN accelerators with better performance and energy efficiency. However, the low-access-cost characteristics of stacked DRAM and the distributed manner of memory access and data storing require us to rebalance the hardware design and DNN mapping. In this paper, we propose NicePIM to efficiently explore the design space of hardware architecture and DNN mapping of DRAM-PIM-based DNN inference accelerators, which consists of three key components: PIM-Tuner, PIM-Mapper and Data-Scheduler. PIM-Tuner optimizes the hardware configurations leveraging a DNN model for classifying area-compliant PIM-node designs and a deep kernel learning model for identifying better hardware parameters. PIM-Mapper explores a variety of DNN mapping configurations, including parallelism between branches of DNN, DNN layer partitioning, DRAM capacity allocation and data layout pattern in DRAM to generate high-hardware-utilization DNN mapping schemes for various hardware configurations. The Data-Scheduler employs an integer-linear-programming-based data scheduling algorithm to alleviate the inter-PIM-node communication overhead of data-sharing brought by DNN layer partitioning. Experimental results demonstrate that NicePIM can optimize hardware configurations for DRAM-PIM systems effectively and can generate high-quality DNN mapping schemes with latency and energy cost reduced by 37% and 28% on average respectively compared to the baseline method.
Junpeng Wang 0002, Mengke Ge, Bo Ding 0004, Qi Xu 0004, Song Chen 0001, Yi Kang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2024 Graph Attention-Based Symmetry Constraint Extraction for Analog Circuits
abstract
In recent years, analog circuits have received extensive attention and are widely used in many emerging applications. The high demand for analog circuits necessitates shorter circuit design cycles. To achieve the desired performance and specifications, various geometrical symmetry constraints must be carefully considered during the analog layout process. However, the manual labeling of these constraints by experienced analog engineers is a laborious and time-consuming process. To handle the costly runtime issue, we propose a graph-based learning framework to automatically extract symmetric constraints in analog circuit layout. The proposed framework leverages the connection characteristics of circuits and the devices’ information to learn the general rules of symmetric constraints, which effectively facilitates the extraction of device-level constraints on circuit netlists. The experimental results demonstrate that compared to state-of-the-art symmetric constraint detection approaches, our framework achieves higher accuracy and F$_1$-score.
Qi Xu 0004, Jing Wang 0131, Lin Cheng 0001, Song Chen 0001, Yi Kang
IEEE Trans. Circuits Syst. I Regul. Pap.1
2024 AD2VNCS: Adversarial Defense and Device Variation-tolerance in Memristive Crossbar-based Neuromorphic Computing Systems
abstract
In recent years, memristive crossbar-based neuromorphic computing systems (NCS) have obtained extremely high performance in neural network acceleration. However, adversarial attacks and conductance variations of memristors bring reliability challenges to NCS design. First, adversarial attacks can fool the neural network and pose a serious threat to security critical applications. However, device variations lead to degradation of the network accuracy. In this article, we propose DFS (Deep neural network Feature importance Sampling) and BFS (Bayesian neural network Feature importance Sampling) training strategies, which consist of Bayesian Neural Network (BNN) prior setting, clustering-based loss function, and feature importance sampling techniques, to simultaneously combat device variation, white-box attack, and black-box attack challenges. Experimental results clearly demonstrate that the proposed training framework can improve the NCS reliability.
Yongtian Bi, Qi Xu 0004, Hao Geng, Song Chen 0001, Yi Kang
ACM Trans. Design Autom. Electr. Syst.2
2024 Floorplanning with Edge-aware Graph Attention Network and Hindsight Experience Replay
abstract
In this article, we focus on chip floorplanning, which aims to determine the location and orientation of circuit macros simultaneously, so the chip area and wirelength are minimized. As the highest level of abstraction in hierarchical physical design, floorplanning bridges the gap between the system-level design and the physical synthesis, whose quality directly influences downstream placement and routing. To tackle chip floorplanning, we propose an end-to-end reinforcement learning (RL) methodology with a hindsight experience replay technique. An edge-aware graph attention network (EAGAT) is developed to effectively encode the macro and connection features of the netlist graph. Moreover, we build a hierarchical decoder architecture mainly consisting of transformer and attention pointer mechanism to output floorplan actions. Since the RL agent automatically extracts knowledge about the solution space, the previously learned policy can be quickly transferred to optimize new unseen netlists. Experimental results demonstrate that, compared with state-of-the-art floorplanners, the proposed end-to-end methodology significantly optimizes area and wirelength on public GSRC and MCNC benchmarks.
Qi Xu 0004, Hao Geng, Song Chen 0001, Bei Yu 0001, Yi Kang
ACM Trans. Design Autom. Electr. Syst.2
2023 Mixed-Type Wafer Failure Pattern Recognition
abstract
The ongoing evolution in process fabrication enables us to step below the 5nm technology node. Although foundries can pattern and etch smaller but more complex circuits on silicon wafers, a multitude of challenges persist. For example, defects on the surface of wafers are inevitable during manufacturing. To increase the yield rate and reduce time-to-market, it is vital to recognize these failures and identify the failure mechanisms of these defects. Recently, applying machine learning-powered methods to combat single defect pattern classification has made significant progress. However, as the processes become increasingly complicated, various single-type defect patterns may emerge and be coupled on a wafer and thus shape a mixed-type pattern. In this paper, we will survey the recent pace of progress on advanced methodologies for wafer failure pattern recognition, especially for mixed-type one. We sincerely hope this literature review can highlight the future directions and promote the advancement of the wafer failure pattern recognition.
Hao Geng, Qi Sun 0002, Tinghuan Chen, Qi Xu 0004, Tsung-Yi Ho, Bei Yu 0001
ASP-DAC4
2023 Tolerating Device-to-Device Variation for Memristive Crossbar-Based Neuromorphic Computing Systems: A New Bayesian Perspective
abstract
Memristive crossbar-based architecture provides an energy-efficient platform to accelerate neural networks (NNs) thanks to its Processing-in-Memory (PIM) nature. However, the device-to-device variation (DDV), which is typically modeled as Lognormal distribution, deviates the programmed weights from their target values, resulting in significant performance degradation. This paper proposes a new Bayesian Neural Network (BNN) approach to enhance the robustness of weights against DDV. Instead of using the widely-used Gaussian variational posterior in conventional BNNs, our approach adopts a DDV-specific variational posterior distribution, i.e., Lognormal distribution. Accordingly, in the new BNN approach, the prior distribution is modified to keep consistent with the posterior distribution to avoid expensive Monte Carlo simulations. Furthermore, the mean of the prior distribution is dynamically adjusted in accordance with the mean of the Lognormal variational posterior distribution for better convergence and accuracy. Compared with the state-of-the-art approaches, experimental results show that the proposed new BNN approach can significantly boost the inference accuracy with the consideration of DDV on several well-known datasets and modern NN architectures. For example, the inference accuracy can be improved from 18% to 74% in the scenario of ResNet-18 on CIFAR-10 even under large variations.
Qi Xu 0004, Bo Yuan 0006
IJCNN2
2023 Reliability-Driven Memristive Crossbar Design in Neuromorphic Computing Systems
abstract
In recent years, memristive crossbar-based neuromorphic computing systems (NCS) have provided a promising solution to the acceleration of neural networks. However, stuck-at faults (SAFs) in the memristor devices significantly degrade the computing accuracy of NCS. Besides, the memristor suffers from the process variations, causing deviation of the actual programming resistance from its target resistance. In this paper, we propose a reliability-driven design framework for a memristive crossbar-based NCS in combination with general and chip-specific design optimizations. First, we design a general reliability-aware training scheme to enhance the robustness of NCS to SAFs and device variations; a dropconnect-inspired approach is developed to alleviate the impact of SAFs; a new weighted error function, including cross-entropy error (CEE), the$l_{2}$-norm of weights, and the sum of squares of first-order derivatives of CEE with respect to weights, is proposed to obtain a smooth error curve, where the effects of variations are suppressed. Second, given the neural network model generated by the reliability-aware training scheme, we exploit chip-specific mapping and re- training to further improve computation accuracy loss incurred by SAFs. Experimental results show that the proposed method can boost the computation accuracy of NCS and improve the NCS robustness. Note to Practitioners—This work is motivated by the manufacturing reliability problem in a memristive crossbar-based NCS. To enhance the robustness of an NCS to SAFs and device variations, this paper presents a reliability-driven design framework with taking account of both general and chip-specific design optimizations. The experimental results have demonstrated that the proposed framework is superior to the prior arts, and can be easily integrated with existing industrial hardware-based fault tolerance solutions for higher accuracy at lower overhead. Memristive crossbar-based computing system gives hope for the anticipated efficient implementation of artificial neuromorphic networks. With the help of the reliability-driven designs, the computation accuracy is restored, and hence we can expect the wide use of memristive crossbar-based computing system in neuromorphic computing applications.
Qi Xu 0004, Junpeng Wang 0002, Bo Yuan 0006, Qi Sun 0002, Song Chen 0001, Bei Yu 0001, Yi Kang, Feng Wu 0001
IEEE Trans Autom. Sci. Eng.1
2023 A Cooperative Multiagent Reinforcement Learning Framework for Droplet Routing in Digital Microfluidic Biochips
abstract
Digital microfluidic biochips (DMFBs) have shown great advantages in automatically executing biochemical protocols through manipulating discrete nano/picoliter droplets which are transported in parallel to achieve high-throughput outcomes. However, because of electrode degradations, the droplet transportation may fail, causing incorrect fluidic operations. To perform safety-critical bio-protocols, the reliability of droplet transportation becomes an utmost concern for DMFBs. It has been shown by the previous works that a reliable transportation policy can be learned using reinforcement learning (RL)-based methods by capturing the underlying health conditions of electrodes and making online decisions. However, previous RL methods may fail to accomplish routing tasks with multiple droplets, because there is a lack of cooperation among different agents (each agent represents one droplet). To deal with this problem and scale RL methods to many droplets, this article proposes a new cooperative centralized learning and distributed execution multiagent RL (MARL) framework for droplet routing in DMFBs using value-decomposition networks (VDNs). Moreover, to speed up the training and decision process as well as apply our method in large biochips, we use a partial observation space where agents can only observe environment in a limited field of view (FOV) centered around themselves. Compared with the state-of-the-art approach, the superior performance of the proposed approach is demonstrated on different DMFBs in terms of success rate and average completion time. We also validate our method on large biochips (e.g.,$\mathbf {50\times 50}$DMFBs) with more droplets than state-of-the-art approach (e.g., ten droplets).
Rong-Quan Yang, Qi Xu 0004, Hailong Yao 0002, Tsung-Yi Ho, Bo Yuan 0006
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2023 Task Modules Partitioning, Scheduling and Floorplanning for Partially Dynamically Reconfigurable Systems with Heterogeneous Resources
abstract
Some field programmable gate arrays (FPGAs) can be partially dynamically reconfigurable with heterogeneous resources distributed on the chip. FPGA-based partially dynamically reconfigurable system (FPGA-PDRS) can be used to accelerate computing and improve computing flexibility. However, the traditional design of FPGA-PDRS is based on manual design. Implementing the automation of FPGA-PDRS needs to solve the problems of task modules partitioning, scheduling, and floorplanning on heterogeneous resources. Existing works only partly solve problems for the automation process of FPGA-PDRS or model homogeneous resources for FPGA-PDRS. To better solve the problems in the automation process of FPGA-PDRS and narrow the gap between algorithm and application, in this paper, we propose a complete workflow including three parts: pre-processing to generate the lists of task module candidate shapes according to the resource requirements, exploration process to search the solution of task modules partitioning, scheduling, and floorplanning, and post-optimization to improve the floorplan success rate. Experimental results show that, compared with state-of-the-art work, the pre-processing process can reduce the occupied area of task modules by 6% on average; the proposed complete workflow can improve performance by 9.6%, and reduce communication cost by 14.2% with improving the resources reuse rate of the heterogeneous resources on the chip. Based on the solution generated by the exploration process, the post-optimization process can improve the floorplan success rate by 11%.
Bo Ding 0004, Jinglei Huang, Junpeng Wang 0002, Qi Xu 0004, Song Chen 0001, Yi Kang
ACM Trans. Design Autom. Electr. Syst.4
2023 Memory-aware Partitioning, Scheduling, and Floorplanning for Partially Dynamically Reconfigurable Systems
abstract
Partially dynamic reconfiguration (PDR) technology can accelerate the reconfiguration process and overcome hardware resource constraints when facing the challenge of high performance with respect to applications and resources constraints on field-programmable gate arrays (FPGAs). On FPGAs with PDR technology, the available on-chip Block RAM (BRAM) resources may not satisfy the memory requirements for all data. If we reserve more BRAM resources, then the total area of the dynamically reconfigurable region (DRR) that is used for calculation will decrease, with a reduction in system performance. We propose a memory-aware optimization framework to search for the optimal solution considering partitioning, scheduling, and floorplanning, where we make a tradeoff between performance and on-chip memory resources utilization. We then propose methods for memory allocation: An ILP model and a heuristic algorithm are provided to determine the minimum memory requirements and the number of corresponding memory blocks for data, as well as to determine whether the memory block with its stored data is assigned on-chip or off-chip by formulating the problem into a 0-1 knapsack problem and solving it using dynamic programming. Experimental results show that the memory-aware optimization framework and methods of memory allocation can increase the amount of on-chip data access to 29.65% of the total data volume with guaranteed performance.
Bo Ding 0004, Jinglei Huang, Qi Xu 0004, Junpeng Wang 0002, Song Chen 0001, Yi Kang
ACM Trans. Design Autom. Electr. Syst.3
2023 DDAM: Data Distribution-Aware Mapping of CNNs on Processing-In-Memory Systems
abstract
Convolution neural networks (CNNs) are widely used algorithms in image processing, natural language processing and many other fields. The large amount of memory access of CNNs is one of the major concerns in CNN accelerator designs that influences the performance and energy-efficiency. With fast and low-cost memory access, Processing-In-Memory (PIM) system is a feasible solution to alleviate the memory concern of CNNs. However, the distributed manner of data storing in PIM systems is in conflict with the large amount of data reuse of CNN layers. Nodes of PIM systems may need to share their data with each other before processing a CNN layer, leading to extra communication overhead. In this article, we propose DDAM to map CNNs onto PIM systems with the communication overhead reduced. Firstly, A data transfer strategy is proposed to deal with the data sharing requirement among PIM nodes by formulating a Traveling-Salesman-Problem (TSP). To improve data locality, a dynamic programming algorithm is proposed to partition the CNN and allocate a number of nodes to each part. Finally, an integer linear programming (ILP)-based mapping algorithm is proposed to map the partitioned CNN onto the PIM system. Experimental results show that compared to the baselines, DDAM can get a higher throughput of 2.0× with the energy cost reduced by 37% on average.
Junpeng Wang 0002, Haitao Du, Bo Ding 0004, Qi Xu 0004, Song Chen 0001, Yi Kang
ACM Trans. Design Autom. Electr. Syst.4
2022 PPATuner: pareto-driven tool parameter auto-tuning in physical design via gaussian process transfer learning
abstract
Thanks to the amazing semiconductor scaling, incredible design complexity makes the synthesis-centric very large-scale integration (VLSI) design flow increasingly rely on electronic design automation (EDA) tools. However, invoking EDA tools especially the physical synthesis tool may require several hours or even days for only one possible parameters combination. Even worse, for a new design, oceans of attempts to navigate high quality-of-results (QoR) after physical synthesis have to be made via multiple tool runs with numerous combinations of tunable tool parameters. Additionally, designers often puzzle over simultaneously considering multiple QoR metrics of interest (e.g., delay, power, and area). To tackle the dilemma within finite resource budget, designing a multi-objective parameter auto-tuning framework of the physical design tool which can learn from historical tool configurations and transfer the associated knowledge to new tasks is in demand. In this paper, we propose PPATuner, a Pareto-driven physical design tool parameter tuning methodology, to achieve a good trade-off among multiple QoR metrics of interest (e.g., power, area, delay) at the physical design stage. By incorporating the transfer Gaussian process (GP) model, it can autonomously learn the transfer knowledge from the existing tool parameter combinations. The experimental results on industrial benchmarks under the 7nm technology node demonstrate the merits of our framework.
Hao Geng, Qi Xu 0004, Tsung-Yi Ho, Bei Yu 0001
DAC2
2022 High-Speed Adder Design Space Exploration via Graph Neural Processes
abstract
Adders are the primary components in the data-path logic of a microprocessor, and thus, adder design has been always a critical issue in the very large-scale integration (VLSI) industry. However, it is infeasible for designers to obtain optimal adder architecture by exhaustively running EDA flow due to the extremely large design space. Previous arts have proposed the machine learning-based framework to explore the design space. Nevertheless, they fall into suboptimality due to a two-stage flow of the learning process and less efficient nor effective feature representations of prefix adder structures. In this article, we first integrate a variational graph autoencoder and a neural process (NP) into an end-to-end, multibranch framework, which is termed thegraph neural process. The former performs automatic feature learning of prefix adder structures, whilst the latter one is designed as an alternative to the Gaussian process. Then, we propose a sequential optimization framework with the graph NP as the surrogate model to explore the Pareto-optimal prefix adder structures with tradeoff among Quality-of-Result (QoR) metrics, such as power, area, and delay. The experimental results show that compared with state-of-the-art methodologies, our framework can achieve a much better Pareto frontier in multiple QoR metric spaces with fewer design-flow evaluations.
Hao Geng, Yuzhe Ma, Qi Xu 0004, Jin Miao, Subhendu Roy, Bei Yu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2022 GoodFloorplan: Graph Convolutional Network and Reinforcement Learning-Based Floorplanning
abstract
Electronic design automation (EDA) comprises a series of computationally difficult optimization problems that require substantial specialized knowledge as well as a considerable amount of trial-and-error efforts. However, open challenges, including long simulation runtime and lack of generalization, continue to restrict the applications of the existing EDA tools. Recently, learning-based algorithms, especially reinforcement learning (RL), have been successfully applied to handle various combinatorial optimization problems by automatically acquiring knowledge from the past experience. In this article, we formulate the floorplanning problem, the first stage of the physical design flow, as a Markov decision process (MDP). An end-to-end learning-based floorplanning framework GoodFloorplan is proposed to explore the design space, which combines graph convolutional network (GCN) and RL. Experimental results demonstrate that compared with state-of-the-art heuristic-based floorplanners, the proposed GoodFloorplan can provide better area and wirelength.
Qi Xu 0004, Hao Geng, Song Chen 0001, Bo Yuan 0006, Cheng Zhuo, Yi Kang, Xiaoqing Wen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2022 Fortune: A New Fault-Tolerance TSV Configuration in Router-Based Redundancy Structure
abstract
In three-dimensional integrated circuits (3D-ICs), through silicon via (TSV) is a critical technique in providing vertical connections. However, yield is one of the key obstacles to adopt the TSV-based 3D-ICs technology in the industry. Various fault-tolerance structures using redundant TSVs to repair faulty functional TSVs have been proposed in the literature for yield and reliability enhancement. However, the TSV repair paths under delay constraint cannot always be generated due to the lack of appropriate repair algorithms. In this article, we propose an effective TSV repair strategy for the router-based TSV redundancy architecture, taking into account the delay overhead. First, we prove that the router-based fault-tolerance structure configuration (RFSC) with the delay constraint is equivalent to the length-bounded multicommodity flow (LBMCF) problem. Then, an integer linear programming (ILP) formulation with acceptable scalability is presented to solve the LBMCF problem. The experimental results demonstrate that, compared with state-of-the-art fault-tolerance designs, the proposed ILP model can provide higher yield and lower delay overhead.
Qi Xu 0004, Hao Geng, Tianming Ni, Song Chen 0001, Bei Yu 0001, Yi Kang, Xiaoqing Wen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2022 Cellular Structure-Based Fault-Tolerance TSV Configuration in 3D-IC
abstract
In 3-D integrated circuits (3D-ICs), through silicon via (TSV) is a critical technique in providing vertical connections. However, the yield is one of the key obstacles to adopt the TSV-based 3D-ICs technology in industry. Various fault-tolerance structures using redundant TSVs to repair faulty functional TSVs have been proposed in literature for yield and reliability enhancement. But the TSV repair paths under delay constraint cannot always be generated due to the lack of appropriate repair algorithms. In this article, we propose an effective TSV repair strategy for the cellular TSV redundancy architecture, with taking account of the delay overhead. First, we prove that the cellular structure-based fault-tolerance TSV configuration with the delay constraint (CSFTC) is equivalent to the length-bounded multicommodity flow (LBMCF) problem. Next, an integer linear programming formulation is presented to solve the LBMCF problem. Finally, to speed-up the fault-tolerance structure configuration process, an efficient Lagrangian relaxation-based heuristic method is further proposed. Experimental results demonstrate that, compared with the state-of-the-art fault-tolerance structures, the proposed method can provide high yield and low delay overhead.
Qi Xu 0004, Song Chen 0001, Yi Kang, Xiaoqing Wen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2022 Synthesizing Brain-network-inspired Interconnections for Large-scale Network-on-chips
abstract
Brain network is a large-scale complex network with scale-free, small-world, and modularity properties, which largely supports this high-efficiency massive system. In this article, we propose to synthesize brain-network-inspired interconnections for large-scale network-on-chips. First, we propose a method to generate brain-network-inspired topologies with limited scale-free and power-law small-world properties, which have a low total link length and extremely low average hop count approximately proportional to the logarithm of the network size. In addition, given the large-scale applications, considering the modularity of the brain-network-inspired topologies, we present an application mapping method, including task mapping and deterministic deadlock-free routing, to minimize the power consumption and hop count. Finally, a cycle-accurate simulator BookSim2 is used to validate the architecture performance with different synthetic traffic patterns and large-scale test cases, including real-world communication networks for the graph processing application. Experiments show that, compared with other topologies and methods, the brain-network-inspired network-on-chips (NoCs) generated by the proposed method present significantly lower average hop count and lower average latency. Especially in graph processing applications with a power-law and tightly coupled inter-core communication, the brain-network-inspired NoC has up to 70% lower average hop count and 75% lower average latency than mesh-based NoCs.
Mengke Ge, Xiaobing Ni, Qi Xu 0004, Song Chen 0001, Jinglei Huang, Yi Kang, Feng Wu 0001
ACM Trans. Design Autom. Electr. Syst.3
2021 Reliability-Driven Neuromorphic Computing Systems Design
abstract
In recent years, memristive crossbar-based neuromorphic computing systems (NCS) have provided a promising solution to the acceleration of neural networks. However, stuck-at faults (SAFs) in the memristor devices significantly degrade the computing accuracy of NCS. Besides, memristors suffer from process variations, causing the deviation of actual programming resistance from its target resistance. In this paper, we propose a novel reliability-driven design framework for a memristive crossbar-based NCS in combination with general and chip-specific design optimizations. First, we design a general reliability-aware training scheme to enhance the robustness of NCS to SAFs and device variations; a dropout-inspired approach is developed to alleviate the impact of SAFs; a new weighted error function, including cross-entropy error (CEE), the l2-norm of weights, and the sum of squares of first-order derivatives of CEE with respect to weights, is proposed to obtain a smooth error curve, where the effects of variations are suppressed. Second, given the neural network model generated by the reliability-aware training scheme, we exploit chip-specific mapping and retraining to further reduce the computation accuracy loss incurred by SAFs. Experimental results clearly demonstrate that the proposed method can boost the computation accuracy of NCS and improve the NCS robustness.
Qi Xu 0004, Junpeng Wang 0002, Hao Geng, Song Chen 0001, Xiaoqing Wen
DATE1
2021 A Cost-Effective TSV Repair Architecture for Clustered Faults in 3-D IC
abstract
Due to the winding level of the thinned wafers and the surface roughness of silicon dies, the through-silicon vias (TSVs) defect tend to be clustered, reducing the yield of 3-D integrated circuit significantly. To tackle this fault clustering problem, the existing TSV repair methods adopt the TSV redundancy idea, which brings a major cost to 3-D integration. In this brief, a honeycomb-TDMA TSV design is proposed to mitigate the impact of multiple clustered faults without the need of redundant TSVs (RTSVs), thereby decreasing the area overhead and enhances the yield. The yield of the honeycomb-TDMA architecture can achieve 91.38%-99.67% for different benchmark circuits from IWLS 2005, which has the highest yield. Furthermore, our design achieves total additional hardware (timing delay overhead) reduction by 83.70%-86.85% (46.01%-55.96%), 66.89%-73.25% (29.41%-38.49%), 68.02%-74.20% (41.40%-52.20%), 60.60%-68.18% (18.09%-33.18%), and 75.86%-80.52% (3.05%-20.91%), respectively, compared with router-based, ring-based, group-based, cellular-based, and honeycomb-based methods. Therefore, the proposed architecture is the best choice in terms of yields, hardware overhead, and timing delay.
Tianming Ni, Qi Xu 0004, Zhengfeng Huang, Huaguo Liang, Aibin Yan, Xiaoqing Wen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2020 Synthesizing A Generalized Brain-inspired Interconnection Network for Large-scale Network-on-chip Systems
abstract
Brain network is a large-scale complex network with scale-free, small-world, and modularity properties, which to a large extent supports this high-efficiency massively parallel computing system known in the world. In this paper, we propose a three-stage method to synthesize a brain-inspired interconnection network for large-scale network-on-chip systems, which minimizes communication hop count, dynamic power consumption, and energy-delay-product. Topology generation, core assignment, and routing path allocation are executed in these three stages, respectively. Experimental results show that our synthesis method can construct large-scale brain-inspired NoC systems with higher communication efficiency and superior performance compared to the state-of-the-art.
Mengke Ge, Qi Xu 0004, Huajie Ruan, Xiaobing Ni, Song Chen 0001, Yi Kang
ACM Great Lakes Symposium on VLSI2
2020 Reliability-Driven Neural Network Training for Memristive Crossbar-Based Neuromorphic Computing Systems
abstract
In recent years, memristive crossbar-based neuromorphic computing systems (NCS) have provided a promising solution to the acceleration of neural networks. However, stuck-at faults (SAFs) in the memristor devices significantly degrade the computing accuracy of NCS. Besides, the memristor suffers from the process variations, causing deviation of the actual programming resistance from its target resistance. In this paper, we propose a reliability-driven network training framework for a memristive crossbar-based NCS, with taking account of both SAFs and device variations challenges. A dropout-inspired approach is first developed to alleviate the impact of SAFs. A new weighted error function, including cross-entropy error (CEE), the l2-norm of weights, and the sum of squares of first-order derivatives of CEE with respect to weights, is further proposed to obtain a smooth error curve, where the effects of variations are suppressed. Experimental results show that the proposed method can boost the computation accuracy of NCS and improve the NCS robustness.
Junpeng Wang 0002, Qi Xu 0004, Bo Yuan 0006, Song Chen 0001, Bei Yu 0001, Feng Wu 0001
ISCAS2
2020 Fault tolerance in memristive crossbar-based neuromorphic computing systems
Qi Xu 0004, Song Chen 0001, Hao Geng, Bo Yuan 0006, Bei Yu 0001, Feng Wu 0001, Zhengfeng Huang
Integr.1
2020 Generalized Fault-Tolerance Topology Generation for Application-Specific Network-on-Chips
abstract
The network-on-chips (NoCs)-based communication architecture is a promising candidate for addressing communication bottlenecks in many-core processors and neural network processors. In this article, we consider the generalized fault-tolerance topology generation problem, where the link (physical channel) or switch failures can happen, for application-specific NoCs (ASNoCs). With a user-defined maximum number of faults, K, we propose an integer linear programming (ILP)-based method to generate ASNoC topologies, which can tolerate at most K faults in switches or links. Given the communication requirements between cores and their floorplan, we first propose a convex-cost-flow-based method to solve a core mapping (CM) problem for building connections between the cores and switches. Second, an ILP-based method is proposed to solve the routing path allocation (PA) problem, where K+1 switch-disjoint routing paths are allocated for every communication flow between the cores. Finally, to reduce switch sizes, we propose to share the switch ports for the connections between the cores and switches and formulate the port sharing problem as a clique-partitioning problem, which is solved by iteratively finding a set of the maximum cliques. Additionally, we propose an ILP-based method to simultaneously solve the CM and routing PA problems when only physical link failures are considered. The experimental results show that the power consumption of fault-tolerance topologies increases almost linearly with K because of the routing path redundancy for fault tolerance. When both switch faults and link faults are considered, port sharing can reduce the average power consumption of fault-tolerance topologies with K = 1, K = 2, and K = 3 by 18.08%, 28.88%, and 34.20%, respectively. When considering only the physical link faults, the experimental results show that compared to the fault-tolerant topology generation (FTTG) algorithm, the proposed method reduces power consumption and hop count by 10.58% and 6.25%, respectively; compared to the de Bruijn Digraph (DBG)-based method, the proposed method reduces power consumption and hop count by 21.72% and 9.35%, respectively.
Song Chen 0001, Mengke Ge, Jinglei Huang, Qi Xu 0004, Feng Wu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2020 Integrated Optimization of Partitioning, Scheduling, and Floorplanning for Partially Dynamically Reconfigurable Systems
abstract
Confronted with the challenge of high performance for applications and the restriction of hardware resources for field-programmable gate arrays (FPGAs), partial dynamic reconfiguration technology is anticipated to accelerate the reconfiguration process and alleviate the device shortage. In this paper, we propose an integrated optimization framework for task partitioning, scheduling, and floorplanning on partially dynamically reconfigurable FPGAs. The partition, schedule, and floorplan of the tasks are represented by the partitioned sequence triple (PST) (PS, QS, RS), where (PS, QS) is a hybrid nested sequence pair for representing the spatial and temporal partitions, as well as the floorplan, and RS is the partitioned dynamic configuration order of the tasks. The floorplanning and scheduling of task modules can be computed from the P-ST in O(n2) time. To integrate the exploration of the scheduling and floorplanning design space, we use a simulated annealing-based search engine and elaborate a perturbation method, where a randomly chosen task module is removed from the partition sequence triple and then reinserted into a proper position selected from all the O(n3) possible combinations of partition, schedule and floorplan. We also prove a sufficient and necessary condition for the feasibility of the partitioning of tasks and scheduling of task configurations, and derive conditions for the feasibility of the insertion points in a P-ST. The experimental results demonstrate the efficiency and effectiveness of the proposed framework.
Song Chen 0001, Jinglei Huang, Bo Ding 0004, Qi Xu 0004
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2020 Memristive Crossbar Mapping for Neuromorphic Computing Systems on 3D IC
abstract
In recent years, neuromorphic computing systems based on memristive crossbar have provided a promising solution to enable acceleration of neural networks. However, most of the neural networks used in realistic applications are often sparse. If such sparse neural network is directly implemented on a single memristive crossbar, then it would result in inefficient hardware realizations. In this work, we propose E3D-FNC, an enhanced three-dimesnional (3D) floorplanning framework for neuromorphic computing systems, in which the neuron clustering and the layer assignment are considered interactively. First, in each iteration, hierarchical clustering partitions neurons into a set of clusters under the guidance of the proposed distance metric. The optimal number of clusters is determined by L-method. Then matrix re-ordering is proposed to re-arrange the columns of the weight matrix in each cluster. As a result, the reordered connection matrix can be easily mapped into a set of crossbars with high utilizations. Next, since the clustering results will in turn affect the floorplan, we perform the floorplanning of neurons and crossbars again. All the proposed methodologies are embedded in an iterative framework to improve the quality of NCS design. Finally, a 3D floorplan of neuromorphic computing systems is generated. Experimental results show that E3D-FNC can achieve highly hardware-efficient designs compared to the state of the art.
Qi Xu 0004, Hao Geng, Song Chen 0001, Bei Yu 0001, Feng Wu 0001
ACM Trans. Design Autom. Electr. Syst.1
2020 Architecture of Cobweb-Based Redundant TSV for Clustered Faults
abstract
In this brief, a cobweb-based redundant through-silicon-via (TSV) design is proposed with efficient hardware as well as high repair rate to repair clustered faulty TSVs (FTSVs). The experimental simulation results demonstrate that for highly clustered faults, the repair rate of the proposed RTSV method is 48.59% and 1.75% higher than that of the ring-based and router-based RTSV methods, respectively. Furthermore, the proposed design can achieve 63.93% and 16.34% hardware reductions compared with the router-based and the ring-based design, respectively.
Tianming Ni, Dongsheng Liu 0001, Qi Xu 0004, Zhengfeng Huang, Huaguo Liang, Aibin Yan
IEEE Trans. Very Large Scale Integr. Syst.3
2019 Adaptive 3D-IC TSV Fault Tolerance Structure Generation
abstract
In 3-D integrated circuits (3D-ICs), through silicon via (TSV) is a critical technique in providing vertical connections. However, the yield is one of the key obstacles to adopt the TSV-based 3D-ICs technology in industry. Various fault-tolerance structures using spare TSVs to repair faulty functional TSVs have been proposed in literature for yield and reliability enhancement, but a valid structure cannot always be found due to the lack of effective generation methods for fault-tolerance structures. In this paper, we focus on the problem of adaptive fault-tolerance structure (AFTS) generation. Given the relations between functional TSVs and spare TSVs, we first calculate the maximum number of tolerant faults in each TSV group. Then we propose an integer linear programming-based model to construct the AFTS with minimal multiplexer delay overhead and hardware cost. We further develop a speed-up technique through an efficient min-cost-max-flow model. All the proposed methodologies are embedded in a top-down TSV planning framework to form functional TSV groups and generate AFTSs. Experimental results show that, compared with state-of-the-art, the number of spare TSVs used for fault tolerance can be effectively reduced.
Song Chen 0001, Qi Xu 0004, Bei Yu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2018 Memristive Crossbar Mapping for Neuromorphic Computing Systems on 3D IC
abstract
In recent years, neuromorphic computing systems based on memristive crossbar have provided a promising solution to enable acceleration of neural networks. Meanwhile, most of the neural networks used in realistic applications are often sparse. If such sparse neural network is directly implemented on a single memristive crossbar, it would result in inefficient hardware realizations. In this work, we propose 3D-FNC, a 3D floorplanning framework for neuromorphic computing systems in consideration of both crossbar utilization and design cost. 3D-FNC groups neurons that connect more common neurons into one cluster, where the optimal number of clusters is determined by L-method. As a result, the connections of a neural network can be effectively mapped to memristive crossbars or discrete synapses. Finally, a 3D floorplanning for memristive crossbars and neurons is developed to reduce area and wirelength cost. Experimental results show that 3D-FNC can achieve highly hardware-efficient designs, compared to state-of-the-art.
Qi Xu 0004, Song Chen 0001, Bei Yu 0001, Feng Wu 0001
ACM Great Lakes Symposium on VLSI1
2017 An Integrated Optimization Framework for Partitioning, Scheduling and Floorplanning on Partially Dynamically Reconfigurable FPGAs
abstract
This paper proposes an integrated optimization framework for task partitioning, scheduling, and floorplanning on partially dynamically reconfigurable FPGAs. In the framework, three problems are represented by a partitioned sequence triple (PS, MS, RS), where (PS, MS) is a hybrid nested sequence pair for floorplanning and RS is a reconfiguration sequence for scheduling. The floorplan and schedule of tasks can be computed from the sequence triple in O(n^2) time. To integrate the exploration of the scheduling and floorplanning design space, a fast perturbation method is elaborated with a simulated annealing-based search engine, where a randomly chosen task is removed from the sequence triple and then inserted back into a proper position selected from all the n^3 possible combinations of partitions, schedule and floorplan. The experimental results demonstrate the efficiency and effectiveness of the proposed framework.
Qi Xu 0004, Jinglei Huang, Song Chen 0001
ACM Great Lakes Symposium on VLSI2
2017 Fast thermal analysis for fixed-outline 3D floorplanning
Qi Xu 0004, Song Chen 0001
Integr.1
2017 Clustered Fault Tolerance TSV Planning for 3-D Integrated Circuits
abstract
In 3-D integrated circuits (3-D ICs), through silicon via (TSV) is a critical technique to provide vertical connections. However, the yield and reliability challenge of TSV in industry is one of key obstacles to adopt the 3-D ICs technology. Various fault-tolerance structures by using additional spare TSVs (s-TSVs) to repair faulty functional TSVs (f-TSVs) have been proposed in literature for yield and reliability enhancement. However, these structures are formed in standard cell placement stage where all the f-TSVs are already placed. In reality, since the s-TSVs can be only inserted into the whitespace, the quality of the generated repair solution is strongly dependent on the whitespace distribution. In this paper, we propose an efficient TSV planning and repair framework in floorplanning stage, which takes nonuniform TSV distribution and clustered TSV defect-distribution into account. The proposed framework mainly consists of four stages: 1) a whitespace redistribution algorithm that uses a probability-based strategy to make the whitespace distribution more reasonable for the f-TSV planning. Subsequently, a convex-cost flow-based model for f-TSV allocation considering the fault clustering; 2) a top-down globally partitioning combined with a bottom-up locally merging to partition f-TSVs into groups with minimum hardware cost; 3) the min-cost max-flow algorithm for s-TSV allocation with minimum wirelength overhead; and 4) an integer linear programming-based model to form a fault-tolerance structure with minimum multiplexer delay overhead. The experimental results demonstrate that the proposed repair framework can improve the yield with minimum hardware cost and multiplexer delay overhead.
Qi Xu 0004, Song Chen 0001, Bei Yu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1