Xiaotao Jia

dblp:154/3032 · DBLP profile ↗
← Back
28ranked-venue papers
8as first author
17since 2021 · last 2026
0000-0003-2207-6092ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 24 · 6 first-author · 13 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Efficient SRAM-PIM Co-Design by Joint Exploration of Value-Level and Bit-Level Sparsity
abstract
Processing-in-memory (PIM) architectures mitigate the Von Neumann bottleneck by integrating computation units into memory arrays. Among PIM architectures, digital SRAMPIM has become a prominent approach, directly integrating digital logic within the SRAM array. However, the rigid crossbar architecture and full array activation pose challenges in efficiently utilizing value-level sparsity. Moreover, neural network models exhibit a high proportion of zero bits within non-zero values, which remain underutilized due to architectural constraints. To overcome these limitations, we present Dyadic Block PIM (DB-PIM), a groundbreaking algorithm-architecture co-design framework to harness both value-level and bit-level sparsity. At the algorithm level, our hybrid-grained pruning technique, combined with a novel sparsity pattern, enables effective sparsity management. Architecturally, DB-PIM incorporates a sparse network and customized digital SRAM-PIM macros, including input pre-processing unit (IPU), dyadic block multiply units (DBMUs), and Canonical Signed Digit (CSD)-based adder trees. It circumvents structured zero values in weights and bypasses unstructured zero bits within non-zero weights and block-wise all-zero bit columns in input features. As a result, the DBPIM framework skips a majority of unnecessary computations, thereby driving significant gains in computational efficiency. Experimental results demonstrate that our DB-PIM framework achieves up to 8.01× speedup and 85.28% energy savings, significantly boosting computational efficiency in digital SRAMPIM systems.
Cenlin Duan, Jianlei Yang 0001, Yiou Wang, Yingjie Qi, Xiaolin He, Bonan Yan, Xiaotao Jia, Weisheng Zhao 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.9
2025 MBQ: Modality-Balanced Quantization for Large Vision-Language Models
abstract
Vision-Language Models (VLMs) have enabled a variety of real-world applications. The large parameter size of VLMs brings large memory and computation overhead which poses significant challenges for deployment. Post-Training Quantization (PTQ) is an effective technique to reduce the memory and computation overhead. Existing PTQ methods mainly focus on large language models (LLMs), without considering the differences across other modalities. In this paper, we discover that there is a significant difference in sensitivity between language and vision tokens in large VLMs. Therefore, treating tokens from different modalities equally, as in existing PTQ methods, may over-emphasize the insensitive modalities, leading to significant accuracy loss. To deal with the above issue, we propose a simple yet effective method, Modality-Balanced Quantization (MBQ), for large VLMs. Specifically, MBQ incorporates the different sensitivities across modalities during the calibration process to minimize the reconstruction loss for better quantization parameters. Extensive experiments show that MBQ can significantly improve task accuracy by up to 4.4% and 11.6% under W3A16 and W4A8 quantization for 7B to 70B VLMs, compared to SOTA baselines. Additionally, we implement a W3A16 GPU kernel that fuses the dequantization and GEMV operators, achieving a 1.4× speedup on LLaVA-onevision-7B on the RTX 4090. The code is available at https://github.com/thu-nics/MBQ.
Yingchun Hu, Xuefei Ning, Xihui Liu, Ke Hong, Xiaotao Jia, Yaqi Yan, Pei Ran, Guohao Dai 0001, Shengen Yan, Huazhong Yang, Yu Wang 0002
CVPR6
2025 MIRACLE: Multimodal Information Retrieval via a Combined In-Memory Processing and Content Addressable Memory Approach
abstract
The rapid advancement of information technology has brought multimodal information retrieval into the research spotlight. Neural networks, particularly Transformers, have emerged as the dominant solution for extracting multimodal feature vectors. While neural network acceleration has been extensively explored, the subsequent retrieval stage in multimodal scenarios remains under-optimized. Conventional retrieval approaches, such as cosine similarity sorting on von Neumann architectures, suffer from significant data migration and computational inefficiencies. Hashing methods enhance storage and computation efficiency but encounter challenges in energy-efficient implementation and mitigating accuracy losses due to modal heterogeneity. This paper presents a hybrid architecture that integrates in-memory processing (PIM) and content-addressable memory (CAM) to address these challenges. Transformer-extracted features are processed via in-memory random hashing leveraging device-intrinsic properties, with CAM facilitating parallel search space reduction. A final cosine similarity reranking stage refines the results while balancing accuracy with energy efficiency. Experimental evaluations validate that the proposed method, when compared to the baseline traditional CPU-based cosine similarity retrieval, 1) achieves almost identical level of accuracy, dramatically outperforming other pure CAMbased Hamming distance retrieval approaches; and 2) reduces latency by $9.45 \times$ and energy consumption by $30.20 \times$.
Xuehui Liu, Tianyang Yu, Shuo Ran, Bi Wu 0002, Xiaotao Jia, Weiqiang Liu 0001, Gang Qu 0001, Weisheng Zhao 0001
DAC7
2025 Ultra Energy-Efficient Butterfly Counting in Bipartite Networks via Algorithm-Architecture Co-Optimization
abstract
Butterfly counting (BFC) problem, which counts the number of butterfly structure in a graph, is fundamental in bipartite network analysis. Recently, considerable efforts toward accelerations of BFC on both CPU and GPU platforms have been reported. However, the underlying BFC algorithms require repetitive vertex traversal and suffer from substantial latency and energy consumption concerns because data in large graphs has very limited reusability. In this paper, we introduce a hardware-software co-optimization approach to tackle these issues. A key innovation behind our approach is an algorithm that employs iterative lightweight arithmetical operations and facilitates highly parallel and pipelined processing. We further develop optimized data compression and pruning strategies to improve the efficiency of processing sparse data. These pivotal advancements are seamlessly integrated with a purpose-built hardware architecture to augment the overall implementation efficiency. Our proposed strategies are thoroughly evaluated on Zynq UltraScale+ FPGA platform. Compared with the state-of-the-art CPU (with 512 GB DRAM) and CPU+GPU (with 128 GB DRAM) implementations, our design achieve speedups of 15.84×and 1.35×, respectively, with only 4 GB DRAM. Meanwhile, our design’s energy efficiency is 50.14× over the CPU+GPU accelerator.
Jianlei Yang 0001, Xiaotao Jia, Gang Qu 0001, Weisheng Zhao 0001
ICCAD4
2025 Finesse: An Agile Design Framework for Pairing-based Cryptography via Software/Hardware Co-Design
abstract
Pairing-based cryptography (PBC) is crucial in modern cryptographic applications.With the rapid advancement of adversarial research and the growing diversity of application requirements, PBC accelerators need regular updates in algorithms, parameter configurations, and hardware design.However, traditional design methodologies face significant challenges, including prolonged design cycles, difficulties in balancing performance and flexibility, and insufficient support for potential architectural exploration.To address these challenges, we introduce Finesse, an agile design framework based on co-design methodology.Finesse leverages a co-optimization cycle driven by a specialized compiler and a multi-granularity hardware simulator, enabling both optimized performance metrics and effective design space exploration.Furthermore, Finesse adopts a modular design flow to significantly shorten design cycles, while its versatile abstraction ensures flexibility across various curve families and hardware architectures.Finesse offers flexibility, efficiency, and rapid prototyping, comparing with previous frameworks.With compilation times reduced to minutes, Finesse enables faster iteration cycles and streamlined hardware-software co-design.Experiments on popular curves * Both authors contributed equally to this research.
Tianwei Pan, Tianao Dai, Jianlei Yang 0001, Hongbin Jing, Zeyu Hao, Xiaotao Jia, Chunming Hu, Weisheng Zhao 0001
ISCA7
2025 AM-CIM: Approximate Memory Based Near Sensor Compute-in-Memory Architecture for Keyword Spotting
abstract
Compute-In-Memory (CIM) has emerged as a promising solution to address the von-Neumann bottleneck, making it a key technology for intelligent computing in edge IoT devices, particularly for real-time applications like keyword spotting (KWS). However, traditional CIM architectures face challenges such as high resource consumption, especially in data conversion, which can significantly impact chip area and energy efficiency. To address these challenges, this work proposes a computational CIM architecture utilizing multilevel analog memory, named AM-CIM, tailored for near-sensor (NS) computation of real-time KWS applications. Additionally, approximate memory technology is integrated into the AM-CIM architecture, employing data resilience scheduling for analog memory which contributes to significant reductions in hardware overhead. This integration facilitates a hardware-software co-design approach. To deploy KWS tasks in AM-CIM, a gated recurrent unit (GRU) network, referred to as MAC-GRU, is implemented. By employing Mel-energy as the input feature at the near-sensor end, the system achieves a 93.13% reduction in feature extraction power consumption. Evaluation results based on TSMC 180-nm technology demonstrate that the AM-CIM architecture achieves an accuracy of 88.51% for 10-keyword classification with a power consumption of$546~\mu W$, while reducing analog memory area by 43.32%.
Xiaotao Jia, Guangcai Yuan, Jianyi Yu, Cong Shi 0003, Qi Wei 0001, Youguang Zhang, Weisheng Zhao 0001, Fei Qiao
IEEE Trans. Circuits Syst. I Regul. Pap.2
2024 Towards Efficient SRAM-PIM Architecture Design by Exploiting Unstructured Bit-Level Sparsity
abstract
Bit-level sparsity in neural network models harbors immense untapped potential. Eliminating redundant calculations of randomly distributed zero-bits significantly boosts computational efficiency. Yet, traditional digital SRAM-PIM architecture, limited by rigid crossbar architecture, struggles to effectively exploit this unstructured sparsity. To address this challenge, we propose Dyadic Block PIM (DB-PIM), a groundbreaking algorithm-architecture co-design framework. First, we propose an algorithm coupled with a distinctive sparsity pattern, termed a dyadic block (DB), that preserves the random distribution of non-zero bits to maintain accuracy while restricting the number of these bits in each weight to improve regularity. Architecturally, we develop a custom PIM macro that includes dyadic block multiplication units (DBMUs) and Canonical Signed Digit (CSD)-based adder trees, specifically tailored for Multiply-Accumulate (MAC) operations. An input pre-processing unit (IPU) further refines performance and efficiency by capitalizing on block-wise input sparsity. Results show that our proposed co-design framework achieves a remarkable speedup of up to 7.69× and energy savings of 83.43%.
Cenlin Duan, Jianlei Yang 0001, Yiou Wang, Yingjie Qi, Xiaolin He, Bonan Yan, Xiaotao Jia, Weisheng Zhao 0001
DAC9
2024 Hierarchical Placement Algorithm for Analog Circuit With Polygonal Modules
abstract
With the increasing application of analog circuits, manual layout processes become time-consuming. Existing research on automatic placement primarily applies to rectangular modules, and it is challenging to simultaneously meet the requirements of speed, compactness, and high performance. This work presents a hierarchical placement algorithm for analog circuit, satisfying various constraints at the device level and distinguishing critical signal path modules from others at a higher level. To minimize the area and the Half Perimeter Wire length (HPWL), while considering the signal path penalty for critical hierarchies, a polygon edges searching method is proposed, which specifically supports polygons and rectangles. The experimental results demonstrate that the automatic placement achieves comparable simulation performance to the manual one, reflecting the swiftness and effectiveness of the proposed method.
Mengzhe Han, Xiaotao Jia, Yingchun Hu
ISCAS2
2024 Miniaturized and Integrated On-Chip Ag/AgCl Micro-electrodes for Chemical Detection
abstract
This work proposes a fabrication and optimization process of miniature Ag/AgCl reference electrodes (RE) at a micrometer (µm) scale for the miniaturized chemical sensors, and these electrodes could also be integrated into a single CMOS chip with solid-state sensors, e.g., ISFET. The thin silver films are first grown by magnetron sputtering and chemically chloridized afterward to form the Ag/AgCl microelectrodes. Different chlorination durations and current densities have been applied to investigate the best fabrication recipe. Both scanning electron microscope (SEM) and energy dispersive spectroscopy (EDS) have been used to analyze the physical composition of electrodes under different recipes. The electrical performance of the proposed micro-electrode has been compared with a standard saturated calomel electrode (SCE), and it is observed that the electrode fabricated with an optimized recipe could provide less than 4 mV accuracy within the pH range from 4 to 10 and more than 200 minutes product life even with only 1.95 × 105µm2area.
Xiaotao Jia, Yuanqi Hu
ISCAS2
2024 LLP-ECCA: A Low-Latency and Programmable Framework for Elliptic Curve Cryptography Accelerators
abstract
Elliptic curve cryptography (ECC) plays a pivotal role in safeguarding data integrity and authentication in contemporary communication contexts, particularly within the domain of Intelligent Transport Systems (ITS). In the realm of ITS, vehicles communicate via the V2X (vehicle-to-everything) protocol, necessitating low-latency responses and minimal power consumption. Given the evolving nature of V2X protocol standards across the globe, programmability becomes a rigid requirement. However, existing strategies cannot meet all these vehicular equipment demands. This paper introduces a novel framework tailored for ECC acceleration to address the issues. Specifically, we propose the design of an Application Specific Instruction Set Processor (ASIP), augmented by pipeline and dual-issue techniques. Furthermore, the envisioned ASIP integrates a hybrid control framework founded on Finite State Machines (FSM), facilitating agile and effective management. Notably, a general GF(p256) Barrett modular multiplier is specially devised to optimize latency and area utilization. Experimental results on Xilinx Kintex Ultrscale+ FPGA demonstrate that the proposed ECC accelerator generates a signature within 131us and verifies a message within 181us, and the performance meets the requirements of today’s V2X standard.
Tianao Dai, Jianlei Yang 0001, Zhaojun Lu, Xiaotao Jia, Gang Qu 0001, Weisheng Zhao 0001
ITC-Asia6
2024 DDC-PIM: Efficient Algorithm/Architecture Co-Design for Doubling Data Capacity of SRAM-Based Processing-in-Memory
abstract
Processing-in-memory (PIM), as a novel computing paradigm, provides significant performance benefits from the aspect of effective data movement reduction. SRAM-based PIM has been demonstrated as one of the most promising candidates due to its endurance and compatibility. However, the integration density of SRAM-based PIM is much lower than other nonvolatile memory-based ones, due to its inherent 6T structure for storing a single bit. Within comparable area constraints, SRAM-based PIM exhibits notably lower capacity. Thus, aiming to unleash its capacity potential, we propose DDC-PIM, an efficient algorithm/architecture co-design methodology that effectively doubles the equivalent data capacity. At the algorithmic level, we propose a filter-wise complementary correlation (FCC) algorithm to obtain a bitwise complementary pair. At the architecture level, we exploit the intrinsic cross-coupled structure of 6T SRAM to store the bitwise complementary pair in their complementary states$(Q/\overline {Q})$, thereby maximizing the data capacity of each SRAM cell. The dual-broadcast input structure and reconfigurable unit support both depthwise and pointwise convolution, adhering to the requirements of various neural networks. Evaluation results show that DDC-PIM yields about$2.84\times $speedup on MobileNetV2 and$2.69\times $on EfficientNet-B0 with negligible accuracy loss compared with PIM baseline implementation. Compared with state-of-the-art SRAM-based PIM macros, DDC-PIM achieves up to$8.41\times $and$2.75\times $improvement in weight density and area efficiency, respectively.
Cenlin Duan, Jianlei Yang 0001, Xiaolin He, Yingjie Qi, Yiou Wang, Ziyan He, Bonan Yan, Xiaotao Jia, Weitao Pan, Weisheng Zhao 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.10
2024 An Energy-Efficient Bayesian Neural Network Implementation Using Stochastic Computing Method
abstract
The robustness of Bayesian neural networks (BNNs) to real-world uncertainties and incompleteness has led to their application in some safety-critical fields. However, evaluating uncertainty during BNN inference requires repeated sampling and feed-forward computing, making them challenging to deploy in low-power or embedded devices. This article proposes the use of stochastic computing (SC) to optimize the hardware performance of BNN inference in terms of energy consumption and hardware utilization. The proposed approach adopts bitstream to represent Gaussian random number and applies it in the inference phase. This allows for the omission of complex transformation computations in the central limit theorem-based Gaussian random number generating (CLT-based GRNG) method and the simplification of multipliers as AND operations. Furthermore, an asynchronous parallel pipeline calculation technique is proposed in computing block to enhance operation speed. Compared with conventional binary radix-based BNN, SC-based BNN (StocBNN) realized by FPGA with 128-bit bitstream consumes much less energy consumption and hardware resources with less than 0.1% accuracy decrease when dealing with MNIST/Fashion-MNIST datasets.
Xiaotao Jia, Huiyi Gu, Jianlei Yang 0001, Weitao Pan, Youguang Zhang, Sorin Cotofana, Weisheng Zhao 0001
IEEE Trans. Neural Networks Learn. Syst.1
2023 NAND-SPIN-based processing-in-MRAM architecture for convolutional neural network acceleration
Yinglin Zhao, Jianlei Yang 0001, Bing Li 0017, Xingzhou Cheng, Xucheng Ye, Xiaotao Jia, Zhaohao Wang, Youguang Zhang, Weisheng Zhao 0001
Sci. China Inf. Sci.7
2023 IMGA: Efficient In-Memory Graph Convolution Network Aggregation With Data Flow Optimizations
abstract
Aggregating features from neighbor vertices is a fundamental operation in graph convolution network (GCN). However, the sparsity in graph data creates poor spatial and temporal locality, causing dynamic and irregular memory access patterns and limiting the performance of aggregation on the Von Neumann architecture. The emerging processing-in-memory (PIM) architecture is based on emerging nonvolatile memory (NVM), like spin-orbit torque magnetic RAM (SOT-MRAM), and demonstrates promising prospects in alleviating the Von Neumann bottleneck. However, the limited memory capacity of PIM medium still incurs non-negligible data movements between PIM architecture and external memory. To solve this challenge, we propose an SOT-MRAM-based in-memory computing architecture, called IMGA, for efficient in-situ graph aggregation. Specifically, we design adaptive data flow management strategies that reuse vertex data in MRAM when processing graphs of different scales and adopt edge data as the control signal source to utilize the graph’s structural information. A reordering optimization strategy leveraging hardware–software co-design principle is proposed to further reduce the costly data movement. Experimental results demonstrate that IMGA achieves an average$2523\times $and$21\times $speedup, and 1.03E+6 and 1.04E+3 energy efficiency compared with CPU and GPU, respectively.
Yuntao Wei, Shangtong Zhang, Jianlei Yang 0001, Xiaotao Jia, Zhaohao Wang, Gang Qu 0001, Weisheng Zhao 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2022 Triangle Counting Accelerations: From Algorithm to In-Memory Computing Architecture
abstract
Triangles are the basic substructure of networks and triangle counting (TC) has been a fundamental graph computing problem in numerous fields such as social network analysis. Nevertheless, like other graph computing problems, due to the high memory-computation ratio and random memory access pattern, TC involves a large amount of data transfers thus suffers from the bandwidth bottleneck in the traditional Von-Neumann architecture. To overcome this challenge, in this paper, we propose to accelerate TC with the emerging processing-in-memory (PIM) architecture through an algorithm-architecture co-optimization manner. To enable the efficient in-memory implementations, we come up to reformulate TC with bitwise logic operations (such as AND), and develop customized graph compression and mapping techniques for efficient data flow management. With the emerging computational Spin-Transfer Torque Magnetic RAM (STT-MRAM) array, which is one of the most promising PIM enabling techniques, the device-to-architecture co-simulation results demonstrate that the proposed TC in-memory accelerator outperforms the state-of-the-art GPU and FPGA accelerations by 12.2x and 31.8x, respectively, and achieves a 34x energy efficiency improvement over the FPGA accelerator.
Jianlei Yang 0001, Yinglin Zhao, Xiaotao Jia, Rong Yin 0001, Xuhang Chen 0001, Gang Qu 0001, Weisheng Zhao 0001
IEEE Trans. Computers4
2022 Accelerating Graph-Connected Component Computation With Emerging Processing-In-Memory Architecture
abstract
Computing the connected component (CC) of a graph is a basic graph computing problem, which has numerous applications like graph partitioning and pattern recognition. Existing methods for computing CC suffer from memory wall problems because of the frequent data transmission between CPU and memory. To overcome this challenge, in this article, we propose to accelerate CC computation with the emerging processing-in-memory (PIM) architecture through an algorithm–architecture co-design manner. The innovation lies in computing CC with bitwise logical operations (such as AND and OR), and the customized data flow management methods to accelerate computation and reduce energy consumption. As a proof of concept, experimental results with computational spin-transfer torque magnetic RAM (STT-MRAM) arrays demonstrate on average$19.8\times $and$12.4\times $speedups compared with the CPU and GPU implementations, and a$35.4 \times $energy efficiency improvement over the CPU implementation. Moreover, we investigate the potential associations between graph computing and bitwise Boolean logic, which could help design more general in-memory graph computing accelerators in the future.
Xuhang Chen 0001, Xiaotao Jia, Jianlei Yang 0001, Gang Qu 0001, Weisheng Zhao 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2021 Efficient Computation Reduction in Bayesian Neural Networks Through Feature Decomposition and Memorization
abstract
The Bayesian method is capable of capturing real-world uncertainties/incompleteness and properly addressing the overfitting issue faced by deep neural networks. In recent years, Bayesian neural networks (BNNs) have drawn tremendous attention to artificial intelligence (AI) researchers and proved to be successful in many applications. However, the required high computation complexity makes BNNs difficult to be deployed in computing systems with a limited power budget. In this article, an efficient BNN inference flow is proposed to reduce the computation cost and then is evaluated using both software and hardware implementations. A feature decomposition and memorization (DM) strategy is utilized to reform the BNN inference flow in a reduced manner. About half of the computations could be eliminated compared with the traditional approach that has been proved by theoretical analysis and software validations. Subsequently, in order to resolve the hardware resource limitations, a memory-friendly computing framework is further deployed to reduce the memory overhead introduced by the DM strategy. Finally, we implement our approach in Verilog and synthesize it with a 45-nm FreePDK technology. Hardware simulation results on multilayer BNNs demonstrate that, when compared with the traditional BNN inference method, it provides an energy consumption reduction of 73% and a 4× speedup at the expense of 14% area overhead.
Xiaotao Jia, Jianlei Yang 0001, Runze Liu 0001, Sorin Cotofana, Weisheng Zhao 0001
IEEE Trans. Neural Networks Learn. Syst.1
2020 TCIM: Triangle Counting Acceleration With Processing-In-MRAM Architecture
abstract
Triangle counting (TC) is a fundamental problem in graph analysis and has found numerous applications, which motivates many TC acceleration solutions in the traditional computing platforms like GPU and FPGA. However, these approaches suffer from the bandwidth bottleneck because TC calculation involves a large amount of data transfers. In this paper, we propose to overcome this challenge by designing a TC accelerator utilizing the emerging processing-in-MRAM (PIM) architecture. The true innovation behind our approach is a novel method to perform TC with bitwise logic operations (such as AND), instead of the traditional approaches such as matrix computations. This enables the efficient in-memory implementations of TC computation, which we demonstrate in this paper with computational Spin-Transfer Torque Magnetic RAM (STT-MRAM) arrays. Furthermore, we develop customized graph slicing and mapping techniques to speed up the computation and reduce the energy consumption. We use a device-to-architecture co-simulation framework to validate our proposed TC accelerator. The results show that our data mapping strategy could reduce 99.99% of the computation and 72% of the memory WRITE operations. Compared with the existing GPU or FPGA accelerators, our in-memory accelerator achieves speedups of 9× and 23.4×, respectively, and a 20.6× energy efficiency improvement over the FPGA accelerator.
Jianlei Yang 0001, Yinglin Zhao, Yingjie Qi, Meichen Liu, Xingzhou Cheng, Xiaotao Jia, Gang Qu 0001, Weisheng Zhao 0001
DAC7
2020 MSFRoute: Multi-Stage FPGA Routing for Timing Division Multiplexing Technique
abstract
As the scale of VLSI circuits and fabrication costs increase rapidly, multi-FPGA prototyping systems are widely adopted in industry to make logic verification faster and cheaper. Since routing signals can usually exceed the number of I/O pins in an FPGA, timing division multiplexing (TDM) technique is required to solve this problem. FPGA routing for developing a prototyping system is a big challenge due to the signal delay of TDM. This paper presents MSFRoute, a multi-stage FPGA routing framework for timing division multiplexing technique, to optimize the signal delay and the routability for prototyping systems. In this work, a TDM ratios assignment algorithm with an efficient parallelization method is proposed to optimize inter-FPGA signal delay. Meanwhile, we propose a practical system clock period optimization method to solve critical signal delay problem. Experimental results show that our routing framework reduces TDM ratios by up to 88.3% with an average reduction rate of 41.8%. With the proposed parallelization method, total flow of MSFRoute can get up to 4.38X speedup with a 2.77X speedup on average.
Zhen Zhuang, Genggeng Liu, Xing Huang 0001, Xiaotao Jia, Wen-Hao Liu 0001, Wenzhong Guo
ACM Great Lakes Symposium on VLSI4
2020 An STT-MRAM based reconfigurable computing-in-memory architecture for general purpose computing
Xiaotao Jia, Jianlei Yang 0001, Weisheng Zhao 0001
CCF Trans. High Perform. Comput.2
2020 Hardware Security in Spin-based Computing-in-memory: Analysis, Exploits, and Mitigation Techniques
abstract
Computing-in-memory (CIM) is proposed to alleviate the processor-memory data transfer bottleneck in traditional von Neumann architectures, and spintronics-based magnetic memory has demonstrated many facilitation in implementing CIM paradigm. Since hardware security has become one of the major concerns in circuit designs, this article, for the first time, investigates spin-based computing-in-memory (SpinCIM) from a security perspective. We focus on two fundamental questions: (1) How can the new SpinCIM computing paradigm be exploited to enhance hardware security?; (2) What security concerns has this new SpinCIM computing paradigm incurred?
Jianlei Yang 0001, Yinglin Zhao, Xiaotao Jia, Gang Qu 0001, Weisheng Zhao 0001
ACM J. Emerg. Technol. Comput. Syst.4
2020 SPINBIS: Spintronics-Based Bayesian Inference System With Stochastic Computing
abstract
Bayesian inference is an effective approach for solving statistical learning problems, especially with uncertainty and incompleteness. However, Bayesian inference is a computing-intensive task whose efficiency is physically limited by the bottlenecks of conventional computing platforms. In this paper, a spintronics-based stochastic computing (SC) approach is proposed for efficient Bayesian inference. The inherent stochastic switching behaviors of spintronic devices are exploited to build a stochastic bitstream generator (SBG) for SC with hybrid CMOS/magnetic tunnel junction (MTJ) circuits design. Aiming to improve the inference efficiency, an SBG sharing strategy is leveraged to reduce the required SBG array scale by integrating a switch network between SBG array and SC logic. A device-to-architecture level framework is proposed to evaluate the performance of spintronics-based Bayesian inference system (SPINBIS). Experimental results on data fusion applications have shown that SPINBIS could improve the energy efficiency about 12× than MTJ-based approach with 45% design area overhead and about 26× than FPGA-based approach.
Xiaotao Jia, Jianlei Yang 0001, Pengcheng Dai, Runze Liu 0001, Yiran Chen 0001, Weisheng Zhao 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2018 Spintronics based stochastic computing for efficient Bayesian inference system
abstract
Bayesian inference is an effective approach for solving statistical learning problems especially with uncertainty and incompleteness. However, inference efficiencies are physically limited by the bottlenecks of conventional computing platforms. In this paper, an emerging Bayesian inference system is proposed by exploiting spintronics based stochastic computing. A stochastic bitstream generator is realized as the kernel components by leveraging the inherent randomness of spintronics devices. The proposed system is evaluated by typical applications of data fusion and Bayesian belief networks. Simulation results indicate that the proposed approach could achieve significant improvement on inference efficiencies in terms of power consumption and inference speed.
Xiaotao Jia, Jianlei Yang 0001, Zhaohao Wang, Yiran Chen 0001, Hai Li 0001, Weisheng Zhao 0001
ASP-DAC1
2018 Electromigration Design Rule aware Global and Detailed Routing Algorithm
abstract
Electromigration (EM) in interconnects is becoming a major concern as the scaling of technology nodes. Electromigration affects chip performance and signal integrity seriously by generating shorts or opens, and then shortens the life-time of integrated circuits. In this paper, we propose an EM-aware routing algorithm in both global and detailed routing stages. Based on physics-based EM modeling and analysis, EM issue is modeled as physical design rule. In global routing stage, an efficient EM-aware Mazerouting algorithm is implemented. An concurrent EM-aware detailed router is then proposed based on multi-commodity flow method. Experimental results show that comparing with general routing algorithm, the proposed EM-aware algorithm could effectively reduce EM risk of signal wires by 92% with slight increasing of wire length and via count.
Xiaotao Jia, Jing Wang 0224, Yici Cai, Qiang Zhou 0001
ACM Great Lakes Symposium on VLSI1
2018 A Multicommodity Flow-Based Detailed Router With Efficient Acceleration Techniques
abstract
Detailed routing is an important stage in very large scale integrated physical design. Due to the extreme scaling of transistor feature size and the complicated design rules, ensuring routing completion without design rule checking (DRC) violations becomes more and more difficult. Studies have shown that the low routing quality partly results from nonoptimal net-ordering nature of traditional sequential methods. The concurrent routing strategy is always based on an NP-hard model, thus is at a disadvantage in runtime. In this paper, we present a novel concurrent detailed routing algorithm that routes all nets simultaneously. Based on the multicommodity flow model, detailed routing problem with complex design rule constraints is formulated as an integer linear programming. Some model simplification heuristics and efficient model solving algorithms are proposed to improve the runtime. Experimental results show that, the proposed algorithms can reduce the DRC violations by 80%, meanwhile can reduce wirelength and via count by 5% and 8% compared with an industry tool. In addition, the proposed algorithm is general that it can be adopted as an incremental detailed router to refine a routing solution, so the number of DRC violations that industry tool cannot fix are further reduced by 27%.
Xiaotao Jia, Yici Cai, Qiang Zhou 0001, Bei Yu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2016 MCFRoute 2.0: A Redundant Via Insertion Enhanced Concurrent Detailed Router
abstract
In modern VLSI design, manufacturing yield and chip performance are seriously affected by via failure. Redundant via insertion is an effective technique recommended by foundries to deal with the via failure. However, due to the extreme scaling of feature size, it is more and more difficult to resolve redundant via insertion (RVI) with limited routing resource while obeying complicated design rules. In this paper, we propose an RVI enhanced concurrent detailed router, MCFRoute 2.0, which effectively avoids design rule violations through a compact integer linear programming (ILP) model. The proposed router can not only route all nets simultaneously but also search for redundant via positions for all via simultaneously during routing stage. In addition, it proposes an RVI aware pin access allocation to further improve the routing performance. Experimental results show that our detailed router outperforms an industry EDA tool that it improves the redundant via insertion rate by 21%, while reducing design rule checking violation count, total wire length and via count by 47%, 4% and 14%, respectively.
Xiaotao Jia, Yici Cai, Qiang Zhou 0001, Bei Yu 0001
ACM Great Lakes Symposium on VLSI1
2016 Secure and Low-Overhead Circuit Obfuscation Technique with Multiplexers
abstract
Circuit obfuscation techniques have been proposed to conceal circuit's functionality in order to thwart reverse engineering (RE) attacks to integrated circuits (IC). We believe that a good obfuscation method should have low design complexity and low performance overhead, yet, causing high RE attack complexity. However, existing obfuscation techniques do not meet all these requirements. In this paper, we propose a polynomial obfuscation scheme which leverages special designed multiplexers (MUXs) to replace judiciously selected logic gates. Candidate to-be-obfuscated logic gates are selected based on a novel gate classification method which utilizes IC topological structure information. We show that this scheme is resilient to all the known attacks, hence it is secure. Experiments are conducted on ISCAS 85/89 and MCNC benchmark suites to evaluate the performance overhead due to obfuscation.
Xiaotao Jia, Qiang Zhou 0001, Yici Cai, Jianlei Yang 0001, Gang Qu 0001
ACM Great Lakes Symposium on VLSI2
2014 MCFRoute: a detailed router based on multi-commodity flow method
abstract
Detailed routing is an important stage in VLSI physical design. Due to the high routing complexity, it is difficult for existing routing methods to guarantee total completion without design rule checking violations (DRCs) and it generally takes several days for designers to fix remaining DRC-s. Studies has shown that the low routing quality partly results from non-optimal net-ordering nature of traditional sequential methods. In this paper, a novel concurrent detailed routing algorithm is presented that overcomes the net-order problem. Based on the multi-commodity flow (M-CF) method, detailed routing problem with complex design rule constraints is formulated as an integer linear programming (ILP) problem. Experiments show that the proposed algorithm is capable of reducing design rule violations while introducing no negative effects on wirelength and via count. Implemented as a detailed router following track assignment, the algorithm can reduce the DRCs by 38%, meantime, wirelength and via count are reduced by 3% and 2.7% respectively comparing with an industry tool. Additionally, the algorithm is adopted as an incremental detailed router to refine a routing solution, and experimental results show that the number of DRCs that industry tool can't fix are further reduce by half. Utilizing the independency between subregions, an efficient parallelization algorithm is implemented that can get a close to linear speedup.
Xiaotao Jia, Yici Cai, Qiang Zhou 0001, Zhuoyuan Li 0003, Zuowei Li
ICCAD1