EDBT 2026 Demo / reviewers in the wild / expert
Jinmei Lai 0001
dblp:50/2813-1 · also Jin-Mei Lai 0001
· DBLP profile ↗
29ranked-venue papers
0as first author
10since 2021 · last 2026
0009-0003-5238-4720ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 28 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mixed-precision Neural Networks on RISC-V CPU with Reconfigurable SIMD Instruction Extension via eFPGAabstractMixed-precision quantization approach has become a key technique for deploying neural networks (NNs) on Central Processing Units (CPUs). However, current CPUs face the following major limitations for mixed-precision NNs: Instruction Set Architecture (ISA) lack adaptability to mixed-precision SIMD instructions, bandwidth bottlenecks in the memory system due to large amounts of vector data, insufficient programmability and flexibility lead to fragile over-optimization as applications change rapidly. In this work, we demonstrate an extended RISC-V CPU with a customized tightly-coupled embedded FPGA (eFPGA) for reconfigurable SIMD instructions extension, targeting mixed-precision NNs deployment. We optimize the vector data mapping strategy with a configurable precision data path for eFPGA access in the CPU extension and design innovative SIMD instructions that extend the RISC-V ISA. We focus on cache hierarchy optimization with related vector register and LSU design choices to achieve high bandwidth SIMD operations. We enable the implementation of high-throughput neural MAC operations at different precisions on a resource-constrained eFPGA fabric via unpacking units and lane-based parallelism. Experimental results demonstrate that our approach, performed on the eFPGA with representative mixed-precision quantized NNs, can achieve an average 23.7× performance and 9.2× energy efficiency gains over the baseline processor. Zixin Yang, Zhichao Wei, Jian Wang 0036, Jinmei Lai 0001 |
ACM Great Lakes Symposium on VLSI | 5 |
| 2025 | An Efficient Traversal Method for FPGA Interconnect Testing Based on Regular Routing
Wenwei Chen, XiaoTong Zhao, Tongshu Ding, Jian Wang 0036, Jinmei Lai 0001 |
FPGA | 6 |
| 2025 | EQViTA: an End-To-End Quantized Vision Transformer Accelerator Implemented on Resource-Constrained FPGAsabstractVision Transformer (ViT) has achieved great success in computer vision tasks, and FPGA-based ViT inference acceleration has recently gained widespread attention. However, the massive number of parameters and intensive matrix computations make it challenging to accelerate ViT models on resource-constrained FPGAs. To address this challenge, prior works have explored ViT quantization and approximate implementations of non-linear operations, but significant hardware resource consumption and potential performance optimization opportunities remain. In this paper, we propose an end-to-end quantized ViT accelerator, EQViTA. Its multi-kernel architecture and time-multiplexed scheduling strategy enable efficient resource utilization and memory-friendly features. We customize designs for the key operators of ViTs. First, we compress the model using INT4 quantization and implement the dataflow of the quantized model through resource-optimized dequantization. Second, we eliminate the convolution hardware module by using the convolution-to-linear mapping, thereby reducing resource usage. Finally, we achieve low-cost and highly parallel acceleration of self-attention and linear transformations through efficient exponential approximation for softmax and an adaptive linear transformation engine. Experiments on the Xilinx ZCU106 FPGA show that compared with state-of-the-art works, EQViTA achieves$1.03 \times$to$8.64 \times$improvements in energy efficiency and$1.14 \times$to$18.6 \times$improvements in normalized throughput. Meanwhile, EQViTA significantly reduces LUT, FF, BRAM, and DSP resource consumption, making it more suitable for resource-constrained FPGAs. Compared to the full-precision DeiT-Tiny model implemented with PyTorch, the INT4 quantized model deployed on EQViTA exhibits a 5.93 % accuracy drop. Jiacheng Cao, Huanlin Luo, Jian Wang 0036, Jinmei Lai 0001 |
FPL | 6 |
| 2024 | Testing Method for Embedded UltraRAM in Field Programmable Gate ArraysabstractMost testing methods for Field Programmable Gate Array (FPGA) on-chip memory are designed for Block RAM, which cannot detect all possible faults in UltraRAM due to its different structures and functions. To efficiently test all potential faults of UltraRAM, a test method with high fault coverage and reduced testing time needs to be designed. In this paper, a complete fault model for UltraRAM testing is established. Existing and new algorithms for testing this UltraRAM fault model are selected or proposed. The algorithms that have been modified or newly designed are tested by fault simulation, and 100% of the sampling faults are detected. The March MSS with DBS can be simplified to reduce the complexity by O(19N), and the complexity of the Byte-wide write enable testing algorithm can be reduced by O(N) when testing port A. By merging configurations to reduce the number of testing configurations, only nine configurations are required to test all the UltraRAMs in the Advanced Micro Devices Versal XCVE2302 device. Jian Wang 0036, Jinmei Lai 0001 |
ATS | 4 |
| 2024 | A Reliable and Efficient Online Solution for Adaptive Voltage and Frequency Scaling on FPGAsabstractAdaptive voltage and frequency scaling (AVFS) technology adjusts the supply voltage and clock frequency based on the actual operating conditions of the circuit. It can significantly improve performance or reduce the power consumption of the device. Existing online field-programmable gate array (FPGA) AVFS solutions have relatively low adjustment efficiency. Many existing solutions rely on offline steps, which do not consider the runtime operating conditions. This article proposes a complete FPGA AVFS solution, which includes a versatile self-checking timing monitor (SCTM) with small resource overhead, efficient AVFS algorithms without any offline steps, and user-friendly comprehensive automation software. Compared with existing online solutions, the proposed solution improves scaling efficiency by reducing the number of configuration times for the clock generation unit. The effectiveness of the solution is evaluated by a set of pubic benchmarks. Experimental results indicate that it can set an appropriate voltage–frequency operating point for the application circuit within dozens of milliseconds. For power-oriented adjustment, the proposed solution can save power ranging from 33.93% to 43.46%, while keeping the frequency not slower than the one reported by the static timing analysis (STA). For performance-oriented adjustment, it can achieve a performance improvement ranging from 60.26% to 101.90% at the nominal voltage. Jiacheng Cao, YaoZhang Liu, Jian Wang 0036, Jinmei Lai 0001, Miaoqing Huang |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2022 | An Effective Test Method for Block RAMs in Heterogeneous FPGAs Based on a Novel Partial Bitstream Relocation TechniqueabstractBlock RAMs (BRAMs) play an important role in modern heterogenous FPGAs, hence how to test them comprehensively and effectively becomes a major concern. On-chip Partial Bitstream Relocation (PBR) technique based on FPGA Dynamic Partial Reconfiguration (DPR) can decrease the time spent on configuring modules in FPGA while reducing the memory resources overhead for storing partial bitstreams of the reconfigurable modules. The previous PBR technique is difficult to be combined with BRAM test directly, because they are somehow tedious, unsuitable for large-scale design or limited to specific devices. Besides, the problem exists for BRAM testing is that fault model is still incomplete and testing algorithms need to be improved to achieve higher fault coverage. An Effective BRAM test method based on a novel PBR technique is proposed in this paper. Our test method establishes a complete fault model for BRAM and improves the testing algorithms for faults in BRAM ECC circuits and intra-word coupling faults in SRAM cells. On-board experiments are carried out with Xilinx xc7vx690t device, and 14 BRAM configurations are used to fully test BRAMs. In conjunction with the proposed PBR technique, the number of configurations can be reduced to 10, which leads to a 35.7% time saving. Changpeng Sun, Huanlin Luo, Jiafeng Liu, Jian Wang 0036, Jinmei Lai 0001, Gang Qu 0001 |
ACM Great Lakes Symposium on VLSI | 7 |
| 2022 | AutoTEA: An Automated Transistor-level Efficient and Accurate design tool for FPGA design
Jiafeng Liu, Jian Wang 0036, Jinmei Lai 0001, Xinxuan Tao, Gang Qu 0001 |
Integr. | 6 |
| 2022 | 3-D Auxiliary Classifier GAN for Hyperspectral Anomaly Detection via Weakly Supervised LearningabstractHyperspectral anomaly detection (AD) is important in Earth observation and remote sensing. However, the low spatial resolution of hyperspectral images, insufficient samples and lack of prior information limit the detection accuracy. To solve these problems, in this paper, we propose an auxiliary classifier generative adversarial network model based on a three-dimensional (3D) convolutional neural network named 3D AC-GAN. Firstly, the model is based on a 3D convolutional neural network design, with 3D tensors as samples. The network maintains valuable image spatial spectrum joint features to achieve good detection results. It can also generate sufficient samples to achieve dataset augmentation, solving the overfitting problem in GAN training. Secondly, we train the model with a weakly supervised method. The label of the samples is obtained through the coarse scanning method. Then, the AC-GAN is trained with the bootstrapping method to mitigate the impact of noise labels. The experimental results show that our proposed algorithm outperforms state-of-the-art AD algorithms. Huanlin Luo, Haowen Zhu, Shengyang Liu, Xinzhong Zhu, Jinmei Lai 0001 |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2021 | AutoTEA: Automated Transistor-level Efficient and Accurate Optimization for GRM FPGA DesignabstractWith the emerging applications such as AI/ML, exploring the FPGA design space for the optimal performance becomes important and also challenging. The popular tool COFFE was built on an academic architecture and cannot be applied directly to modern FPGA chips with GRM (general routing matrix) architecture. In this work, we present our recently developed fully Automated Transistor-level Efficient and Accurate tool, AutoTEA, which features accurate area and delay models, and a fast solution space exploration method for GRM FPGA circuit optimization. The results show that AutoTEA is able to improve a previously manually optimized design (on the tape-out FPGA chip) by 11%. Jiafeng Liu, Jian Wang 0036, Jinmei Lai 0001, Gang Qu 0001 |
FCCM | 5 |
| 2021 | TCP-Net: Minimizing Operation Counts of Binarized Neural Network InferenceabstractBinarized neural network (BNN) dataflow inference accelerators have emerged as a promising solution to be applied in cost- and power-restricted domains, such as loT and smart edge- devices. However, there still exists abundant redundancies in BNN inference, which severely limit the performances of these accelerators. To alleviate the performance degradation, we propose TCP-Net, an efficient architecture to minimize the number of operations in BNN inference while maintaining the original accuracy. Inspired by the observation that the processes of obtaining the outputs of multiple related kernels in BNNs contain significant repeated calculations, we first build a formula to bridge these outputs by utilizing the kernel inclusion similarity and eliminate the unnecessary operations. Through the recursive algorithm, we further convert each original XNOR-popcount convolution into threshold-comparable-popcount (TCP) operations, which can be implemented to directly achieve the final output without any extra steps. Furthermore, we reduce the remaining TCP operation counts by exploiting tile-based pruning strategy. Compared to the prior state-of-the-art designs, TCP-Net saves 79.11 percent of the operations without any accuracy loss, bringing 6.0× inference-speedup and 12.4× energy-efficiency improvement. Qingliang Liu 0002, Jinmei Lai 0001 |
ISCAS | 3 |
| 2020 | HRAE: Hardware-assisted Randomization against Adversarial Example AttacksabstractWith the rapid advancements of the artificial intelligence, machine learning, especially neural networks, have shown huge superiority over humans in image recognition, autonomous vehicles and medical diagnosis. However, its opacity and inexplicability provide many chances for malicious attackers. Recent researches have shown that neural networks are vulnerable to adversarial example (AE) attacks. In the testing stage, it fools the model by adding subtle perturbations to the original sample to misclassify the input, which poses a serious threat to safety-critical areas such as autonomous driving. In order to mitigate this threat, this paper proposes a hardware-assisted randomization method against AEs, where an approximate computing technique in hardware, voltage over-scaling (VOS), is used to randomize the training set of the model, then the processed data are used to generate multiple neural network models, finally multiple redundant models are used for the integrated classification and detection of the AEs. Various AE attacks on the proposed defense are evaluated to prove its effectiveness. Jiliang Zhang 0002, Shuang Peng 0010, Yupeng Hu 0004, Wei Hu 0008, Jinmei Lai 0001, Jing Ye 0001, Xiangqi Wang |
ATS | 6 |
| 2020 | INTB: A New FPGA Interconnect Model for Architecture ExplorationabstractCAD exploration is important for designing FPGA interconnect topologies. It includes two steps: first, design a model with some parameters that can express as much architecture space. Second, use CAD flow to analyze the described interconnect architecture. In this paper, we present a new interconnect model, named INTB (Interconnect Block). At a logical position, one INTB is adopted to represent all related routing resources and hierarchical parameters are designed to simplify description. Compared with existing CB-SB model, INTB model can support more interconnect features of modern FPGA, such as various types of wire segment and complex connections. These features can improve FPGA routing ability. For the application of INTB model, two modifications are made in CAD flow: one is generation of routing resource graph (RRG). A tile-based method is proposed to generate RRG from parameters. The other is cost computing during routing process. Two strategies are applied respectively for cost estimation of short and curve wire segment, which do not exist in CB-SB model. INTB model and CAD improvement are implemented in VTR 8.0. The experiments consist of two parts. First, INTB model is adopted to re-describe CB-SB architectures to verify its description capacity. After CAD flow, average difference of routing area and timing between two models is about 4% and 5%. Second, INTB model is used to explore architecture space with modern FPGA features. Experimental results show obvious performance enhancement, over 10% in some benchmarks. Qinghua Duan, Jian Wang 0036, Jinmei Lai 0001 |
FPGA | 6 |
| 2020 | FPTLOPT: An Automatic Transistor-Level Optimization Tool for GRM FPGAabstractThe FPGA circuit design usually adopts full-custom design method, it indicates that it is difficult to design and optimize an FPGA manually. So, we present FPTLOPT (FPGA Transistor-Level Optimization Tool) which supports a more complex FPGA architecture called general routing matrix (GRM) architecture, and also has higher-accuracy and higher-speed than COFFE [1]. To fit a more complex FPGA architecture, we use the regular matching method to automatically extract the circuits type and build the circuits netlist; To get the higher-accuracy, we predict the layout area by area model we build, then we precisely predict the layout post simulation delay by load model we build; To get the higher-speed, we devise the variable range greedy algorithm, to expanding range automatically. We also provide equalization kernel multi-thread acceleration that can change the thread number according to the current CPU hardware environment. The experimental results illustrate that FPTLOPT supports the optimization of GRM architecture and build the key sub-circuit netlist. Also, the area prediction is by maximum of 43%, the delay get from delay prediction is 28% more precise than the ones in COFFE. Besides, quickly gets the optimal transistor sizing results for different optimization objectives. For the same circuit, the optimization speed is 19.96 times faster than COFFE. Zhengjie Li, Jian Wang 0036, Jinmei Lai 0001 |
FPGA | 4 |
| 2020 | A Tile-based Interconnect Model for FPGA Architecture ExplorationabstractModern FPGA has complex interconnect, like curve wires and two-level local muxs (global wires -> block input pins). Existing interconnect model (CB-SB) cannot describe these routing fabrics, hindering CAD exploration of modern FPGA. Qinghua Duan, Jian Wang 0036, Jinmei Lai 0001 |
ACM Great Lakes Symposium on VLSI | 6 |
| 2019 | Transistor-Level Optimization Methodology for GRM FPGA Interconnect CircuitsabstractDue to its dominance in the whole chip area, power and delay, the FPGA interconnect circuits are traditionally designed by full custom design method. We present an automated transistor-level sizing optimization methodology for GRM FPGA interconnect circuits. In order to get accurate and effective predicated area, the commonly used diffusion sharing, transistor folding and inputs sharing are considered. To get the accurate and effective delay value, we avoid the inaccuracy of using linear device model, and use two schemes to build wire model: the wire within a circuit and the wire between interconnect circuits. To decrease simulation time, we propose multi-thread acceleration method and the Minimum-Final-Delay (MFD) algorithm which optimizes interconnect circuit as a whole, not separated part. For switch box optimization, MFD algorithm requires 38% less number of simulations than COFFE's algorithm. We use 65nm CMOS process technology for evaluation. For different optimization strategy, we emphasize either representative critical path delay or overall layout area. Compare to full-custom design method, the global cost can be decrease by 3% ~ 17%. For different transistor sizing combinations, 10/50 threads can be ~ 9X/15X faster than single-thread. Compared with the manual design method, our optimization methodology explores larger design space, and it decreases the circuit design optimization time from months to hours. Zhengjie Li, Yuanlong Xiao, Yunbing Pang, Jian Wang 0036, Jinmei Lai 0001 |
FPGA | 6 |
| 2019 | An Analytical-based Hybrid Algorithm for FPGA PlacementabstractAs the capacity of FPGA increases, FPGA placers that adopt Simulated Annealing (SA) algorithm take more and more runtime. To solve this problem, this paper presents HCAS, a Hybrid algorithm Combining Analytical method and SA. There are three modifications in HCAS: (1) In global placement, faster and better result is realized by modified analytical algorithm. (2) In detailed placement, proper tradeoff is made between quality and runtime through improvement of SA. (3) Optimization workload of timing and wirelength is reasonably assigned between global and detailed placement according to algorithm features. HCAS is implemented in the newest VPR. Compared to VPR placer, it obtains a speedup of 11.1x, with 3% shorter wirelength and 5% smaller critical path delay. Compared to other analytical-based hybrid placers, HCAS achieves greater speedup and enhancement of placement quality is similar or better. Qinghua Duan, Liran Hu, Zhengjie Li, Meng Yang 0013, Jian Wang 0036, Jinmei Lai 0001 |
ACM Great Lakes Symposium on VLSI | 8 |
| 2019 | An Automatic Transistor-Level Tool for GRM FPGA Interconnect Circuits OptimizationabstractDue to its dominance in FPGA area and delay, the interconnect circuit is traditionally designed and optimized in full customized fashion, which can be extremely time consuming. In this paper, we propose an automated transistor-level sizing optimization method for the widely-used General Routing Matrix FPGA interconnect circuits with the following three features: (1) an area model that takes into account the commonly used diffusion sharing, transistor folding and inputs sharing techniques in order to have an accurate area predication; (2) an accurate and effective non-linear delay model that treats the wire within a circuit and the wire between interconnect circuits separately; (3) a multi-thread acceleration method and the Minimum-Final-Delay algorithm to speed-up the simulation. The global optimization cost is measured by the product of the interconnect circuit area and the representative path delay based on our proposed models. The cost reduces 10.9%, when we use 65nm CMOS process chip for evaluation. The simulation time for different transistor sizing combinations is improved by 9X and 15X when 10 and 50 threads are used, respectively, faster than single-thread. Compared with the manual design method, our proposed optimization approach explores a larger design space and reduces the optimization time from months to hours. Zhengjie Li, Yuanlong Xiao, Yunbing Pang, Jian Wang 0036, Jinmei Lai 0001 |
ACM Great Lakes Symposium on VLSI | 7 |
| 2016 | Testing FPGA Local Interconnects Based on Repeatable Configuration Modules (Abstract Only)abstractThis paper provides a novel technique for testing FPGA local interconnects based on repeatable configuration modules (RCMs). In order to fully detect all the possible faults, local interconnects together with the adjacent logic blocks in an FPGA are programmed to form a set of RCMs that are repeatable all over the FPGA array. After the RCMs for configurable logic blocks (CLBs) and other types of embedded cores (such as digital signal processor, block random access memory) are constructed, test configurations are generated by connecting the RCMs one by one throughout the whole FPGA array. The number of test configurations depends on the structure of the FPGA and the exact types of hard cores inside the FPGA. Experimental results show that a total of 47 test configurations are sufficient to achieve 96.2% fault coverage for Xilinx XC4VLX200 FPGA local interconnects. This project is supported by the State Key Laboratory of ASIC and System, Fudan University, No. 2015MS007. Jian Wang 0036, Meng Yang 0013, Jinmei Lai 0001 |
FPGA | 4 |
| 2014 | Novel FPGA clock network with low latency and skew (abstract only)abstractClock network is a dedicated network for distributing multiple clock signals to every logic modules in a system. Be significantly different from ASIC where the clock tree is custom built by users, clock network in FPGA is usually fixed after chip fabrication and cannot be changed for different user circuits. This paper is committed to design and implement FPGA clock network with low latency and skew. We first propose a novel clock network for FPG, which is a backbone-branches topology and can be easily integrated to the tiled FPGA with reasonable area. There are one clock backbone and several primary clock branches in the network. When the chip scales up, this clock network can be extended easily. Afterwards, series of strategies such as hybrid multiplexer, bypassing, looping back and Programmable Delay Adjustment Unit (DAU) are employed to optimize latency and skew. Moreover, the prominent couple capacitance and crosstalk effect of clock routing in nanometer are also given consideration in physical implementation. This clock network is applied to own-designed FPGA with 65nm technology. Post-layout simulation results indicate that our clock network with normal loads can uphold 600MHz clock with the maximum clock latency and skew being typically 2.22ns and 40ps respectively, 1.79ns and 39ps in the fast case, achieving up to 78.2% improvement for skew as well as 47.5% for latency, compared to a commercial 65nm FPGA device. Jian Wang 0036, Jinmei Lai 0001 |
FPGA | 3 |
| 2013 | FPGA bitstream compression and decompression using LZ and golomb coding (abstract only)abstractIn this paper we propose an optimized bitstream compression algorithm based on LZ and a novel architecture of decompressor, the proposed algorithm improves the Compression Ratio by fully utilizing the regularity of configuration bits of CLB (Configurable Logic Box) in FPGA and using the variable length Golomb coding method. The experimental results show that the Optimized method can improve the Compression Ratio of LZSS by 32.3% for bitstream with high regularity and 10.3% for bitstream with low regularity, and our approach shows a higher flexibility than the BMC+RLE arithmetic when compressing the bitstream with high regularity for various FPGA. Moreover, we design a two-buffer-window decompressor to download the compressed bitstreams. In order to increase the throughput of the proposed decompressor, we design a multi-stage data selector in it. The post-simulation of the decompressor shows that its throughput is up to 9280 Mbps under 65nm CMOS process. And that is 4352Mbps when verified on a Virtex-5 FPGA. Jinsong Mao, Hao Zhou 0008, Haijiang Ye, Jinmei Lai 0001 |
FPGA | 4 |
| 2013 | A novel multithread routing method for FPGAs (abstract only)abstractWe propose a platform-independent multithread routing method for FPGAs including two aspects: single high fanout net is routed parallel within itself and several low fanout nets are routed parallel between themselves. Routing for high fanout nets usually takes considerable time because of the large physical area surrounded by bounding boxes to traverse and tens of terminals to connect. Therefore, one high fanout net is partitioned into several subnets with fewer terminals and smaller bounding boxes to be routed in parallel. However, low fanout nets with intrinsic small bounding boxes and few terminals could hardly be divided. Instead, low fanout nets whose bounding boxes are not overlapping with each other are routed concurrently. A new graph, named bounding box graph, was utilized to facilitate the process of selecting several nets to be routed concurrently. In this graph, one vertex stands for a corresponding net and one edge between two connected vertex means that the two represented nets have their bounding boxes overlapped. Several strategies are introduced to balance the load among threads and ensure the deterministic results. The routing times scale down with increasing number of threads. On a 4-core processor, this technique improves the run-time by ~1.9 × with routing quality degrading by no more than 2.3%. Qiuli Li, Jian Wang 0036, Jinmei Lai 0001 |
FPGA | 4 |
| 2013 | Yet Another Many-Objective Clustering (YAMO-Pack) for FPGA CADabstractIn this paper, Yet Another Many-Objective Clustering (YAMO-Pack) is proposed for academic field programmable gate array architecture model. The YAMO-Pack introduces the impact of attraction between Basic Logic Elements (BLE) and the selected Configurable Logic Blocks (CLB) when BLEs are indirectly connected to the CLBs. Consequently more external nets are absorbed into clusters and speed of the design is improved. Experimental results show that the proposed algorithm outperforms previous approaches in terms of number of CLBs, number of external nets and critical path. Meng Yang 0013, Jinmei Lai 0001, Jiarong Tong |
FPL | 2 |
| 2013 | A novel net-partition-based multithread FPGA routing methodabstractA platform-independent multithread routing method for FPGAs is proposed in this paper. Specifically, the proposed method includes two aspects for maximal parallelization. First, for high fanout net which usually takes considerable time to be routed due to large bounding boxes and number of terminals, it is partitioned into several subnets to be routed in parallel. Second, low fanout nets with non-overlapping bounding boxes are identified and routed in parallel as well to further speed up the routing process. A bounding box graph was constructed to facilitate the process of selecting nets to be routed concurrently. In addition, load balancing and synchronization strategies are introduced to raise routing efficiency and ensure the deterministic results. Experiments on different platforms and benchmarks with various combinations of high and low fanout nets are carried out. This technique improves the run-time by ~1.9 × with routing quality degrading by no more than 2.3%, on a quad-core processor platform. Jian Wang 0036, Jinmei Lai 0001 |
FPL | 3 |
| 2012 | A novel full coverage test method for CLBs in FPGA (abstract only)abstractFPGA's configurability makes it difficult for FPGA's manufacturers to fully test it. In this paper, a full coverage test method for FPGA's Configurable Logic Blocks (CLBs) is proposed, through which all basic logics of FPGA's every CLB can be fully tested. Innovative test circuits are configured to build repeatable logic arrays for look-up tables, distributed random access memories, configurable registers and other logics. The programmable interconnects needed to connect CLBs in these test circuits are also repeatable, making the configuration process much easier and the test speed much faster. The test method is implemented on different scales of Xilinx Virtex chips, where 19 test configuration circuits are needed to achieve 100% coverage for all CLBs. Besides, the method is transplantable and independent of FPGA's array size. To evaluate the test method reliably and guide the process of test vectors generation, a fault simulator - Turbofault is used to simulate FPGA's test coverage. Liguang Chen, Jinmei Lai 0001 |
FPGA | 4 |
| 2009 | A delay-optimized universal FPGA routing architectureabstractA universal FPGA routing architecture is presented, which ensures that every module in the FPGA including CLBs and IOBs have a uniform interconnect architecture, and the load of interconnect lines is equally distributed. So, this architecture is highly repeatable and the signal delay is predictable and regular. Furthermore, the realization of the programmable interconnect point (PIP) and the buffer driver is also optimized to benefit the signal delay up to 5%.The test results of the example chip show the reasonableness of these ideas. Huowen Zhang, Lei Duan, Jinmei Lai 0001, Jiarong Tong |
ASP-DAC | 4 |
| 2009 | A novel minloop SB design to improve FPGA routabilityabstractIn this paper, we present a new design of a switch box of an FPGA that requires less channel width than conventional designs in routing the same circuit. The design, called a Minloop switch box, is based on the method of minimum-loop-size maximization in routing resources. Experimental results show that the Minloop switch box requires 17.7%, 8.0% and 2.4% less channel width than the classic fabric of Disjoint, Universal and Wilton switch boxes, respectively. JIanDe Yu, Jinmei Lai 0001 |
FPGA | 2 |
| 2005 | A design of high speed double precision floating point adder using macro modulesabstractBased on SMIC 0.18 μm 1.8v six-layer-metal CMOS process, we implement a 64-bit high speed pipelined floating point adder which satisfied IEEE 754 standard. After the critical path analysis of the pipelined structure, we custom design three macro modules in order to reduce critical path delay. After placement in datapath style and routing, we implement the layout of floating point adder. The chip area is 1.44 mm2 and clock frequency is 518MHz. Chi Huang, Jinmei Lai 0001, Chengshou Sun |
ASP-DAC | 3 |
| 2005 | Design of A 2.4-GHz integrated frequency synthesizerabstractA 2.4-GHz integrated frequency synthesizer of PLL-based in 0.35-μm RF process is presented. A fully integrated cross-coupled LC VCO of low phase noise is implemented. Prescaler accompanied with phase-switching is used to eliminate the glitch. The charge-pump having excellent current matching performance and wide output voltage range is achieved. The synthesizer has a frequency tuning range from 2.28 to 2.75 GHz. The simulation results show that it dissipates less than 66mW; settle time is less 100us; the phase noise is -117dBc/[email protected] Jinmei Lai 0001, Chengshou Sun |
ASP-DAC | 4 |
| 2003 | Periodic steady-state analysis of coupled ODE-AE-CGE systems for MOS RF autonomous circuit simulationabstractThis paper studies the steady-state analysis of a system of ODE-AE-CGE arising from simulation of MOS RF autonomous circuits in which the transistor plays as current generator and its current equations could adequately represent the transistor large-signal dynamic behavior. Gauss-Seidel relaxation is applied to decouple the system so that shooting method, which is used to find the periodic steady-state response, can be performed with low-order sensitivity matrix. This simplifies the solution process and results in a fast algorithm. We illustrate the new algorithm with the simulation of a typically voltage-controlled oscillator (VCO), the simulation results show good agreement with those obtained by MEDICI, a device simulator. Zaiman Chen, Jinmei Lai 0001, Qianling Zhang, Omar Wing, Junyan Ren |
ASP-DAC | 3 |