Jian Wang 0036

dblp:39/449-36 · DBLP profile ↗
← Back
19ranked-venue papers
0as first author
8since 2021 · last 2026
0000-0002-6086-1924ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 19 · 8 since 2021
YearPublicationVenuePosition
2026 Mixed-precision Neural Networks on RISC-V CPU with Reconfigurable SIMD Instruction Extension via eFPGA
abstract
Mixed-precision quantization approach has become a key technique for deploying neural networks (NNs) on Central Processing Units (CPUs). However, current CPUs face the following major limitations for mixed-precision NNs: Instruction Set Architecture (ISA) lack adaptability to mixed-precision SIMD instructions, bandwidth bottlenecks in the memory system due to large amounts of vector data, insufficient programmability and flexibility lead to fragile over-optimization as applications change rapidly. In this work, we demonstrate an extended RISC-V CPU with a customized tightly-coupled embedded FPGA (eFPGA) for reconfigurable SIMD instructions extension, targeting mixed-precision NNs deployment. We optimize the vector data mapping strategy with a configurable precision data path for eFPGA access in the CPU extension and design innovative SIMD instructions that extend the RISC-V ISA. We focus on cache hierarchy optimization with related vector register and LSU design choices to achieve high bandwidth SIMD operations. We enable the implementation of high-throughput neural MAC operations at different precisions on a resource-constrained eFPGA fabric via unpacking units and lane-based parallelism. Experimental results demonstrate that our approach, performed on the eFPGA with representative mixed-precision quantized NNs, can achieve an average 23.7× performance and 9.2× energy efficiency gains over the baseline processor.
Zixin Yang, Zhichao Wei, Jian Wang 0036, Jinmei Lai 0001
ACM Great Lakes Symposium on VLSI4
2025 An Efficient Traversal Method for FPGA Interconnect Testing Based on Regular Routing
Wenwei Chen, XiaoTong Zhao, Tongshu Ding, Jian Wang 0036, Jinmei Lai 0001
FPGA5
2025 EQViTA: an End-To-End Quantized Vision Transformer Accelerator Implemented on Resource-Constrained FPGAs
abstract
Vision Transformer (ViT) has achieved great success in computer vision tasks, and FPGA-based ViT inference acceleration has recently gained widespread attention. However, the massive number of parameters and intensive matrix computations make it challenging to accelerate ViT models on resource-constrained FPGAs. To address this challenge, prior works have explored ViT quantization and approximate implementations of non-linear operations, but significant hardware resource consumption and potential performance optimization opportunities remain. In this paper, we propose an end-to-end quantized ViT accelerator, EQViTA. Its multi-kernel architecture and time-multiplexed scheduling strategy enable efficient resource utilization and memory-friendly features. We customize designs for the key operators of ViTs. First, we compress the model using INT4 quantization and implement the dataflow of the quantized model through resource-optimized dequantization. Second, we eliminate the convolution hardware module by using the convolution-to-linear mapping, thereby reducing resource usage. Finally, we achieve low-cost and highly parallel acceleration of self-attention and linear transformations through efficient exponential approximation for softmax and an adaptive linear transformation engine. Experiments on the Xilinx ZCU106 FPGA show that compared with state-of-the-art works, EQViTA achieves$1.03 \times$to$8.64 \times$improvements in energy efficiency and$1.14 \times$to$18.6 \times$improvements in normalized throughput. Meanwhile, EQViTA significantly reduces LUT, FF, BRAM, and DSP resource consumption, making it more suitable for resource-constrained FPGAs. Compared to the full-precision DeiT-Tiny model implemented with PyTorch, the INT4 quantized model deployed on EQViTA exhibits a 5.93 % accuracy drop.
Jiacheng Cao, Huanlin Luo, Jian Wang 0036, Jinmei Lai 0001
FPL5
2024 Testing Method for Embedded UltraRAM in Field Programmable Gate Arrays
abstract
Most testing methods for Field Programmable Gate Array (FPGA) on-chip memory are designed for Block RAM, which cannot detect all possible faults in UltraRAM due to its different structures and functions. To efficiently test all potential faults of UltraRAM, a test method with high fault coverage and reduced testing time needs to be designed. In this paper, a complete fault model for UltraRAM testing is established. Existing and new algorithms for testing this UltraRAM fault model are selected or proposed. The algorithms that have been modified or newly designed are tested by fault simulation, and 100% of the sampling faults are detected. The March MSS with DBS can be simplified to reduce the complexity by O(19N), and the complexity of the Byte-wide write enable testing algorithm can be reduced by O(N) when testing port A. By merging configurations to reduce the number of testing configurations, only nine configurations are required to test all the UltraRAMs in the Advanced Micro Devices Versal XCVE2302 device.
Jian Wang 0036, Jinmei Lai 0001
ATS3
2024 A Reliable and Efficient Online Solution for Adaptive Voltage and Frequency Scaling on FPGAs
abstract
Adaptive voltage and frequency scaling (AVFS) technology adjusts the supply voltage and clock frequency based on the actual operating conditions of the circuit. It can significantly improve performance or reduce the power consumption of the device. Existing online field-programmable gate array (FPGA) AVFS solutions have relatively low adjustment efficiency. Many existing solutions rely on offline steps, which do not consider the runtime operating conditions. This article proposes a complete FPGA AVFS solution, which includes a versatile self-checking timing monitor (SCTM) with small resource overhead, efficient AVFS algorithms without any offline steps, and user-friendly comprehensive automation software. Compared with existing online solutions, the proposed solution improves scaling efficiency by reducing the number of configuration times for the clock generation unit. The effectiveness of the solution is evaluated by a set of pubic benchmarks. Experimental results indicate that it can set an appropriate voltage–frequency operating point for the application circuit within dozens of milliseconds. For power-oriented adjustment, the proposed solution can save power ranging from 33.93% to 43.46%, while keeping the frequency not slower than the one reported by the static timing analysis (STA). For performance-oriented adjustment, it can achieve a performance improvement ranging from 60.26% to 101.90% at the nominal voltage.
Jiacheng Cao, YaoZhang Liu, Jian Wang 0036, Jinmei Lai 0001, Miaoqing Huang
IEEE Trans. Very Large Scale Integr. Syst.4
2022 An Effective Test Method for Block RAMs in Heterogeneous FPGAs Based on a Novel Partial Bitstream Relocation Technique
abstract
Block RAMs (BRAMs) play an important role in modern heterogenous FPGAs, hence how to test them comprehensively and effectively becomes a major concern. On-chip Partial Bitstream Relocation (PBR) technique based on FPGA Dynamic Partial Reconfiguration (DPR) can decrease the time spent on configuring modules in FPGA while reducing the memory resources overhead for storing partial bitstreams of the reconfigurable modules. The previous PBR technique is difficult to be combined with BRAM test directly, because they are somehow tedious, unsuitable for large-scale design or limited to specific devices. Besides, the problem exists for BRAM testing is that fault model is still incomplete and testing algorithms need to be improved to achieve higher fault coverage. An Effective BRAM test method based on a novel PBR technique is proposed in this paper. Our test method establishes a complete fault model for BRAM and improves the testing algorithms for faults in BRAM ECC circuits and intra-word coupling faults in SRAM cells. On-board experiments are carried out with Xilinx xc7vx690t device, and 14 BRAM configurations are used to fully test BRAMs. In conjunction with the proposed PBR technique, the number of configurations can be reduced to 10, which leads to a 35.7% time saving.
Changpeng Sun, Huanlin Luo, Jiafeng Liu, Jian Wang 0036, Jinmei Lai 0001, Gang Qu 0001
ACM Great Lakes Symposium on VLSI6
2022 AutoTEA: An Automated Transistor-level Efficient and Accurate design tool for FPGA design
Jiafeng Liu, Jian Wang 0036, Jinmei Lai 0001, Xinxuan Tao, Gang Qu 0001
Integr.5
2021 AutoTEA: Automated Transistor-level Efficient and Accurate Optimization for GRM FPGA Design
abstract
With the emerging applications such as AI/ML, exploring the FPGA design space for the optimal performance becomes important and also challenging. The popular tool COFFE was built on an academic architecture and cannot be applied directly to modern FPGA chips with GRM (general routing matrix) architecture. In this work, we present our recently developed fully Automated Transistor-level Efficient and Accurate tool, AutoTEA, which features accurate area and delay models, and a fast solution space exploration method for GRM FPGA circuit optimization. The results show that AutoTEA is able to improve a previously manually optimized design (on the tape-out FPGA chip) by 11%.
Jiafeng Liu, Jian Wang 0036, Jinmei Lai 0001, Gang Qu 0001
FCCM4
2020 INTB: A New FPGA Interconnect Model for Architecture Exploration
abstract
CAD exploration is important for designing FPGA interconnect topologies. It includes two steps: first, design a model with some parameters that can express as much architecture space. Second, use CAD flow to analyze the described interconnect architecture. In this paper, we present a new interconnect model, named INTB (Interconnect Block). At a logical position, one INTB is adopted to represent all related routing resources and hierarchical parameters are designed to simplify description. Compared with existing CB-SB model, INTB model can support more interconnect features of modern FPGA, such as various types of wire segment and complex connections. These features can improve FPGA routing ability. For the application of INTB model, two modifications are made in CAD flow: one is generation of routing resource graph (RRG). A tile-based method is proposed to generate RRG from parameters. The other is cost computing during routing process. Two strategies are applied respectively for cost estimation of short and curve wire segment, which do not exist in CB-SB model. INTB model and CAD improvement are implemented in VTR 8.0. The experiments consist of two parts. First, INTB model is adopted to re-describe CB-SB architectures to verify its description capacity. After CAD flow, average difference of routing area and timing between two models is about 4% and 5%. Second, INTB model is used to explore architecture space with modern FPGA features. Experimental results show obvious performance enhancement, over 10% in some benchmarks.
Qinghua Duan, Jian Wang 0036, Jinmei Lai 0001
FPGA5
2020 FPTLOPT: An Automatic Transistor-Level Optimization Tool for GRM FPGA
abstract
The FPGA circuit design usually adopts full-custom design method, it indicates that it is difficult to design and optimize an FPGA manually. So, we present FPTLOPT (FPGA Transistor-Level Optimization Tool) which supports a more complex FPGA architecture called general routing matrix (GRM) architecture, and also has higher-accuracy and higher-speed than COFFE [1]. To fit a more complex FPGA architecture, we use the regular matching method to automatically extract the circuits type and build the circuits netlist; To get the higher-accuracy, we predict the layout area by area model we build, then we precisely predict the layout post simulation delay by load model we build; To get the higher-speed, we devise the variable range greedy algorithm, to expanding range automatically. We also provide equalization kernel multi-thread acceleration that can change the thread number according to the current CPU hardware environment. The experimental results illustrate that FPTLOPT supports the optimization of GRM architecture and build the key sub-circuit netlist. Also, the area prediction is by maximum of 43%, the delay get from delay prediction is 28% more precise than the ones in COFFE. Besides, quickly gets the optimal transistor sizing results for different optimization objectives. For the same circuit, the optimization speed is 19.96 times faster than COFFE.
Zhengjie Li, Jian Wang 0036, Jinmei Lai 0001
FPGA3
2020 A Tile-based Interconnect Model for FPGA Architecture Exploration
abstract
Modern FPGA has complex interconnect, like curve wires and two-level local muxs (global wires -> block input pins). Existing interconnect model (CB-SB) cannot describe these routing fabrics, hindering CAD exploration of modern FPGA.
Qinghua Duan, Jian Wang 0036, Jinmei Lai 0001
ACM Great Lakes Symposium on VLSI5
2019 Transistor-Level Optimization Methodology for GRM FPGA Interconnect Circuits
abstract
Due to its dominance in the whole chip area, power and delay, the FPGA interconnect circuits are traditionally designed by full custom design method. We present an automated transistor-level sizing optimization methodology for GRM FPGA interconnect circuits. In order to get accurate and effective predicated area, the commonly used diffusion sharing, transistor folding and inputs sharing are considered. To get the accurate and effective delay value, we avoid the inaccuracy of using linear device model, and use two schemes to build wire model: the wire within a circuit and the wire between interconnect circuits. To decrease simulation time, we propose multi-thread acceleration method and the Minimum-Final-Delay (MFD) algorithm which optimizes interconnect circuit as a whole, not separated part. For switch box optimization, MFD algorithm requires 38% less number of simulations than COFFE's algorithm. We use 65nm CMOS process technology for evaluation. For different optimization strategy, we emphasize either representative critical path delay or overall layout area. Compare to full-custom design method, the global cost can be decrease by 3% ~ 17%. For different transistor sizing combinations, 10/50 threads can be ~ 9X/15X faster than single-thread. Compared with the manual design method, our optimization methodology explores larger design space, and it decreases the circuit design optimization time from months to hours.
Zhengjie Li, Yuanlong Xiao, Yunbing Pang, Jian Wang 0036, Jinmei Lai 0001
FPGA5
2019 An Analytical-based Hybrid Algorithm for FPGA Placement
abstract
As the capacity of FPGA increases, FPGA placers that adopt Simulated Annealing (SA) algorithm take more and more runtime. To solve this problem, this paper presents HCAS, a Hybrid algorithm Combining Analytical method and SA. There are three modifications in HCAS: (1) In global placement, faster and better result is realized by modified analytical algorithm. (2) In detailed placement, proper tradeoff is made between quality and runtime through improvement of SA. (3) Optimization workload of timing and wirelength is reasonably assigned between global and detailed placement according to algorithm features. HCAS is implemented in the newest VPR. Compared to VPR placer, it obtains a speedup of 11.1x, with 3% shorter wirelength and 5% smaller critical path delay. Compared to other analytical-based hybrid placers, HCAS achieves greater speedup and enhancement of placement quality is similar or better.
Qinghua Duan, Liran Hu, Zhengjie Li, Meng Yang 0013, Jian Wang 0036, Jinmei Lai 0001
ACM Great Lakes Symposium on VLSI7
2019 An Automatic Transistor-Level Tool for GRM FPGA Interconnect Circuits Optimization
abstract
Due to its dominance in FPGA area and delay, the interconnect circuit is traditionally designed and optimized in full customized fashion, which can be extremely time consuming. In this paper, we propose an automated transistor-level sizing optimization method for the widely-used General Routing Matrix FPGA interconnect circuits with the following three features: (1) an area model that takes into account the commonly used diffusion sharing, transistor folding and inputs sharing techniques in order to have an accurate area predication; (2) an accurate and effective non-linear delay model that treats the wire within a circuit and the wire between interconnect circuits separately; (3) a multi-thread acceleration method and the Minimum-Final-Delay algorithm to speed-up the simulation. The global optimization cost is measured by the product of the interconnect circuit area and the representative path delay based on our proposed models. The cost reduces 10.9%, when we use 65nm CMOS process chip for evaluation. The simulation time for different transistor sizing combinations is improved by 9X and 15X when 10 and 50 threads are used, respectively, faster than single-thread. Compared with the manual design method, our proposed optimization approach explores a larger design space and reduces the optimization time from months to hours.
Zhengjie Li, Yuanlong Xiao, Yunbing Pang, Jian Wang 0036, Jinmei Lai 0001
ACM Great Lakes Symposium on VLSI6
2016 Testing FPGA Local Interconnects Based on Repeatable Configuration Modules (Abstract Only)
abstract
This paper provides a novel technique for testing FPGA local interconnects based on repeatable configuration modules (RCMs). In order to fully detect all the possible faults, local interconnects together with the adjacent logic blocks in an FPGA are programmed to form a set of RCMs that are repeatable all over the FPGA array. After the RCMs for configurable logic blocks (CLBs) and other types of embedded cores (such as digital signal processor, block random access memory) are constructed, test configurations are generated by connecting the RCMs one by one throughout the whole FPGA array. The number of test configurations depends on the structure of the FPGA and the exact types of hard cores inside the FPGA. Experimental results show that a total of 47 test configurations are sufficient to achieve 96.2% fault coverage for Xilinx XC4VLX200 FPGA local interconnects. This project is supported by the State Key Laboratory of ASIC and System, Fudan University, No. 2015MS007.
Jian Wang 0036, Meng Yang 0013, Jinmei Lai 0001
FPGA2
2014 Novel FPGA clock network with low latency and skew (abstract only)
abstract
Clock network is a dedicated network for distributing multiple clock signals to every logic modules in a system. Be significantly different from ASIC where the clock tree is custom built by users, clock network in FPGA is usually fixed after chip fabrication and cannot be changed for different user circuits. This paper is committed to design and implement FPGA clock network with low latency and skew. We first propose a novel clock network for FPG, which is a backbone-branches topology and can be easily integrated to the tiled FPGA with reasonable area. There are one clock backbone and several primary clock branches in the network. When the chip scales up, this clock network can be extended easily. Afterwards, series of strategies such as hybrid multiplexer, bypassing, looping back and Programmable Delay Adjustment Unit (DAU) are employed to optimize latency and skew. Moreover, the prominent couple capacitance and crosstalk effect of clock routing in nanometer are also given consideration in physical implementation. This clock network is applied to own-designed FPGA with 65nm technology. Post-layout simulation results indicate that our clock network with normal loads can uphold 600MHz clock with the maximum clock latency and skew being typically 2.22ns and 40ps respectively, 1.79ns and 39ps in the fast case, achieving up to 78.2% improvement for skew as well as 47.5% for latency, compared to a commercial 65nm FPGA device.
Jian Wang 0036, Jinmei Lai 0001
FPGA2
2014 A FPGA prototype design emphasis on low power technique
abstract
In this paper, we propose a fully-functional Nanometer FPGA prototype chip. Compared to traditional single supply voltage, single threshold voltage design, we explore low power nanometer FPGA design challenges with Multi-Vt, Static Voltage Scaling and sleep mode technique. Compared to Dynamic Voltage Scaling (DVS), we make a table of Voltage-Delay parameter pairs under different voltage conditions so that timing information can be calculated by a Static Timing Analysis (STA) tool. Thus a lowest supply power is chosen among all results which meet the timing requirements. This approach would simplify the hardware design since we don't need a complex workload detection circuit compared to DVS system. By separating supply voltages, we can directly shutdown power supply of the unused circuits. Compared to inserting sleep transistor in pull-up or pull-down networks, we can eliminate the speed penalty cased by the additional sleep transistor. We implement a tile-based heterogeneous architecture with island style routing and embedded specific blocks such as DSP and memory. The array size is 64×31 (Row×Col) including 64×24 CLBs. The final design is fabricated using a 1P10M 65-nm bulk CMOS process. Test results show a 53% reduction in static power compared to a commercial FPGA device which is also fabricated in 65nm process and has a similar array size.
Jian Wang 0036, Meilai Jin
FPGA2
2013 A novel multithread routing method for FPGAs (abstract only)
abstract
We propose a platform-independent multithread routing method for FPGAs including two aspects: single high fanout net is routed parallel within itself and several low fanout nets are routed parallel between themselves. Routing for high fanout nets usually takes considerable time because of the large physical area surrounded by bounding boxes to traverse and tens of terminals to connect. Therefore, one high fanout net is partitioned into several subnets with fewer terminals and smaller bounding boxes to be routed in parallel. However, low fanout nets with intrinsic small bounding boxes and few terminals could hardly be divided. Instead, low fanout nets whose bounding boxes are not overlapping with each other are routed concurrently. A new graph, named bounding box graph, was utilized to facilitate the process of selecting several nets to be routed concurrently. In this graph, one vertex stands for a corresponding net and one edge between two connected vertex means that the two represented nets have their bounding boxes overlapped. Several strategies are introduced to balance the load among threads and ensure the deterministic results. The routing times scale down with increasing number of threads. On a 4-core processor, this technique improves the run-time by ~1.9 × with routing quality degrading by no more than 2.3%.
Qiuli Li, Jian Wang 0036, Jinmei Lai 0001
FPGA3
2013 A novel net-partition-based multithread FPGA routing method
abstract
A platform-independent multithread routing method for FPGAs is proposed in this paper. Specifically, the proposed method includes two aspects for maximal parallelization. First, for high fanout net which usually takes considerable time to be routed due to large bounding boxes and number of terminals, it is partitioned into several subnets to be routed in parallel. Second, low fanout nets with non-overlapping bounding boxes are identified and routed in parallel as well to further speed up the routing process. A bounding box graph was constructed to facilitate the process of selecting nets to be routed concurrently. In addition, load balancing and synchronization strategies are introduced to raise routing efficiency and ensure the deterministic results. Experiments on different platforms and benchmarks with various combinations of high and low fanout nets are carried out. This technique improves the run-time by ~1.9 × with routing quality degrading by no more than 2.3%, on a quad-core processor platform.
Jian Wang 0036, Jinmei Lai 0001
FPL2