EDBT 2026 Demo / reviewers in the wild / expert
Xi Jin 0002
dblp:74/5217-2
· DBLP profile ↗
30ranked-venue papers
0as first author
15since 2021 · last 2026
0000-0002-4159-2925ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 28 · 14 since 2021Software engineering, systems software and programming languages · 2Computer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient Top-k Sorter through Runtime Lane Selection of Priority Queues on a Directed Ring
Huawen Liang, Wei Yuan 0006, Qizhe Wu, Xi Jin 0002 |
ISCAS | 4 |
| 2025 | Exploring the Performance Improvement of Tensor Processing Engines through Transformation in the Bit-weight Dimension of MACsabstractGeneral matrix-matrix multiplication (GEMM), serving as a cornerstone of AI computations, has positioned tensor processing engines (TPEs) as increasingly critical components within existing GPUs and domain-specific architectures (DSA). Our analysis identifies that the prevailing architectures primarily focus on dataflow or operand reuse strategies, when considering the combination of matrix multiplication with multiply-accumulator (MAC) itself, it provides greater optimization space for the design of TPEs. This work introduces a novel perspective on matrix multiplication from a hardware standpoint, focusing on the bit-weight dimension of MACs. Through this lens, we propose a finer-grained TPE notation, using matrix triple loops as an example, introducing new methods and ideas for designing and optimizing PE microarchitecture. Based on the new notation and transformations, we propose four optimization techniques that achieve varying degrees of improvement in timing, area, and power consumption. We implement our design in RTL using the SMIC-28nm process. Applying our methods to four classic TPE architectures (include systolic array [20], 3D-Cube [27], multiplier-adder tree [48], and 2D-Matrix [30]), we achieved area efficiency improvements of $1.27 \times, 1.28 \times, 1.56 \times$, and $1.44 \times$, and $1.04 \times, 1.56 \times, 1.49 \times$, and $1.20 \times$ for energy efficiency respectively. When applied to a bit-slice architecture, we achieved a $12.10 \times$ improvement in energy efficiency and $2.85 \times$ in area efficiency compared to Laconic [38]. Our Verilog HDL code, along with timing, area, and power reports for circuit synthesis in URL: https://github.com/wqzustc/High-Performance-Tensor-Processing-Engines. Qizhe Wu, Huawen Liang, Yuchen Gui, Zhichen Zeng 0002, Zerong He, Linfeng Tao, Letian Zhao, Zhaoxi Zeng, Wei Yuan 0006, Xi Jin 0002 |
HPCA | 12 |
| 2025 | SageSC: Accelerating GraphSAGE Minibatch Inference on Memory-Intensive GraphsabstractGraph neural networks demonstrate excellent performance on node classification tasks in graph datasets. For inference tasks on memory-intensive graphs, the storage burden, memory access bottlenecks, and load imbalance issues arise. The minibatch inference proposed in GraphSAGE is an effective method for minimizing these problems. However, minibatch inference introduces new challenges: while it facilitates subsequent computation, the irregular random memory access pressure shifts to the minibatch construction phase, creating performance bottlenecks in the system. In this work, to address the aforementioned challenges, we propose a novel scattered minibatch construction and aggregation (SMCA) algorithm to optimize sampling, batch construction, and aggregation computations for minimizing their latency. This method distributes memoryintensive workloads and exploits the parallelism between memory groups. Evaluation results show that the proposed accelerator SageSC achieves speedups ranging from 3x to 96x compared to CPU/GPU baselines, especially on memory-intensive graphs, while outperforming existing state-of-the-art designs. Yuchen Gui, Wei Yuan 0006, Qizhe Wu, Huawen Liang, Letian Zhao, Linfeng Tao, Zhongguang Xu, Xi Jin 0002 |
ICCD | 8 |
| 2025 | MHE-TPE: Multi-Operand High-Radix Encoder for Mixed-Precision Fixed-Point Tensor Processing Engines
Qizhe Wu, Jinyi Zhou, Zhanhe Hu, Zhichen Zeng 0002, Huawen Liang, Jiuru Zhu, Linfeng Tao, Xin Zhang 0176, Zekang Cheng, Letian Zhao, Wei Yuan 0006, Xi Jin 0002 |
MICRO | 13 |
| 2025 | GHVSA: Graph-based high-dimensional vector search accelerator
Wei Yuan 0006, Huawen Liang, Xi Jin 0002 |
J. Syst. Archit. | 3 |
| 2025 | FANNS: An FPGA-Based Approximate Nearest-Neighbor Search AcceleratorabstractApproximate nearest-neighbor search (ANNS) based on high-dimensional vectors has been extensively utilized in data science and neural networks. However, deploying ANNS in production systems requires minimal redundant computation, high recall rates, and low on-chip memory usage, which existing hardware accelerators fail to offer. We propose FANNS, a solution for ANNS based on high-dimensional vectors that can eliminate redundant computations and reuse on-chip data. Extensive evaluations show that FANNS achieves an average of$184.1\times $,$33.0\times $,$2.9\times $, and$2.5\times $better energy efficiency than CPUs, GPUs, and two state-of-the-art ANNS architectures, i.e., DF-GAS and Vstore, respectively. Wei Yuan 0006, Xi Jin 0002 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2024 | A FPGA-HBM-Based Hardware Streaming Accelerator for GNN SamplingabstractSampling takes a long time during GNN training and inference, especially in large-scale graph datasets, so accelerating the sampling process is of great value. Present GNN samplers are mainly focused on the CPU or GPU side, with fewer sampling accelerator implementations on hardware platforms such as FPGAs, and most of the existing accelerators focus on the entire GNN computation process rather than on sampling. Sampling GNNs on the FPGA side is difficult because of the intensive random memory access. In this work, we propose a FPGA-HBM-based streaming sampler working in the time domain to accelerate traditional FPGA sampling, which realizes GNN sampling while reading data in bus bursts and can satisfy both playback and non-playback sampling modes. In addition, the fast loading of node feature vectors is also realized by combining the features of HBM. The proposed hardware achieves a speedup of$2\times$to$20\times$relative to traditional FPGA-based node index sampling and$5\times$to$18\times$relative to feature vector loading at the CPU side. Yuchen Gui, Qizhe Wu, Wei Yuan 0006, Huawen Liang, Xi Jin 0002 |
ASAP | 6 |
| 2024 | RingTK: A Ring, Parallel and High Performance Top-K Sorter on FPGAabstractGetting the K largest/smallest elements from$N$inputs is one of the essential operations in many applications. In this article, we propose RingTK, a ring, parallel, and high performance Top-K sorter implemented on FPGA. We use a priority queue as the basic processing unit and design a Top-K sorter based on a ring topology with a global maximum module(GMM) to obtain good scalability and high performance. Based on the ring topology, we design Ring Multiplexers (RMUX) and modify the GMM to enable RingTK to efficiently handle different K-sizes and concurrent tasks. We design a encoder to flexibly deal with different data formats and max/min Top-K tasks. Finally, we implement the proposed architecture on the Xilinx XCVU37P FPGA. The results show that the proposed architecture has good scalability, and the throughput and parallelism are approximately linear. We can achieve a throughput of 35.38GB/s with 32 PQs that exceed existing literature with K=8160. Huawen Liang, Qizhe Wu, Wei Yuan 0006, Teng Tian, Xi Jin 0002 |
FCCM | 5 |
| 2024 | Efficient Message Passing Architecture for GCN Training on HBM-based FPGAs with Orthogonal Topology On-Chip NetworksabstractGraph Convolutional Networks (GCNs) are state-of-the-art deep learning models for representation learning on graphs. However, the efficient training of GCNs is hampered by constraints in memory capacity and bandwidth, compounded by the irregular data flow that results in communication bottlenecks. To address these challenges, we propose a message-passing architecture that leverages NUMA-based memory access properties and employs a parallel multicast routing algorithm based on a 4-D hypercube network within the accelerator for efficient message passing in graphs. Additionally, we have re-engineered the backpropagation algorithm specific to GCNs within our proposed accelerator. This redesign strategically mitigates the memory demands prevalent during the training phase and diminishes the computational overhead associated with the transposition of extensive matrices. Compared to the state-of-the-art HP-GNN architecture we achieved a performance improvement of 1.03×~1.81×. Qizhe Wu, Letian Zhao, Yuchen Gui, Huawen Liang, Xi Jin 0002 |
FPGA | 6 |
| 2024 | EN-T: Optimizing Tensor Computing Engines Performance via Encoder-Based MethodologyabstractTensor computations, with matrix multiplication being the primary operation, serve as the fundamental basis for data analysis, physics, machine learning, and deep learning. As the scale and complexity of data continue to grow rapidly, the demand for tensor computations has also increased significantly. To meet this demand, several research institutions have started developing dedicated hardware for tensor computations. To further improve the computational performance of tensor process units, we have reexamined the issue of computation reuse that was previously overlooked in existing architectures. As a result, we propose a novel EN-T architecture that can reduce chip area and power consumption. Furthermore, our method is compatible with existing tensor processing units. We evaluated our method on prevalent microarchitectures, the results demonstrate an average improvement in area efficiency of 8.7 %, 12.2 %, and 11.0 % for tensor computing units at computational scales of 256 GOPS, 1 TOPS, and 4 TOPS, respectively. Similarly, there were energy efficiency enhancements of 13.0 %, 17.5 %, and 15.5 %. Qizhe Wu, Yuchen Gui, Zhichen Zeng 0002, Huawen Liang, Xi Jin 0002 |
ICCD | 6 |
| 2023 | MCANet: Multiscale Cross-Modality Attention Network for Multispectral Pedestrian Detection
Letian Zhao, Xi Jin 0002 |
MMM (1) | 4 |
| 2022 | FP-GNN: Adaptive FPGA accelerator for Graph Neural Networks
Teng Tian, Letian Zhao, Qizhe Wu, Wei Yuan 0006, Xi Jin 0002 |
Future Gener. Comput. Syst. | 6 |
| 2022 | G-NMP: Accelerating Graph Neural Networks with DIMM-based Near-Memory Processing
Teng Tian, Letian Zhao, Xuecang Zhang, Fangmin Lu, Xi Jin 0002 |
J. Syst. Archit. | 8 |
| 2022 | QEGCN: An FPGA-based accelerator for quantized GCNs with edge-level parallelism
Wei Yuan 0006, Teng Tian, Qizhe Wu, Xi Jin 0002 |
J. Syst. Archit. | 4 |
| 2021 | A Gather Accelerator for GNNs on FPGA PlatformabstractGraph Neural Networks (GNNs) have emerged as the state-of-the-art deep learning model for representation learning on graphs. GNNs mainly include two phases with different execution patterns. The Gather phase, depends on the structure of the graph, presenting a sparse and irregular execution pattern. The Apply phase, acts like other neural networks, showing a dense and regular execution pattern. It is challenging to accelerate GNNs, due to irregular data communication to gather information within the graph. To address this challenge, hardware acceleration for Gather phase is critical. The purpose of this research is to design and implement an FPGA-based accelerator for Gather phase. It achieves excellent performance on acceleration and energy efficiency. Evaluation is performed using a Xilinx VCU128 FPGA with three commonly-used datasets. Compared to the state-of-the-art software framework running on Intel Xeon CPU and NVIDIA P100 GPU, our work achieves on average 101.28× speedup with 75.27× dynamic energy reduction and average 12.27× speedup with 45.56× dynamic energy reduction, respectively. Wei Yuan 0006, Teng Tian, Huawen Liang, Xi Jin 0002 |
ICPADS | 4 |
| 2020 | Exploration of Memory Access Optimization for FPGA-based 3D CNN AcceleratorabstractThree-dimensional convolutional networks (3D CNNs) are used efficiently in various video recognition applications. Compared to traditional 2D CNNs, extra temporal dimension causes 3D CNNs more computationally intensive and to have a larger memory footprint. Therefore, the memory optimization is extremely crucial in this case. This paper presents a design space exploration of memory access optimization for FPGA-based 3D CNN accelerator. We present a non-overlapping data tiling method for contiguous off-chip memory access and explore on-chip data reuse opportunity by leveraging different loop ordering strategies. We propose a hardware architecture design which can flexibly support different loop ordering strategies for each 3D CNN layer. With the help of hardware/software co-design, we can provide the optimal configuration toward an energy-efficient and high-performance accelerator design. According to the experiments on AlexNet, VGG16, and C3D, our optimal model reduces up to 84% DRAM accesses and 55% energy consumption on C3D compared to a baseline model, and demonstrates state-of-the-art performance compared to prior FPGA implementations. Teng Tian, Xi Jin 0002, Letian Zhao |
DATE | 2 |
| 2020 | FPDeep: Scalable Acceleration of CNN Training on Deeply-Pipelined FPGA ClustersabstractDeep convolutional Neural Networks (CNNs) have revolutionized numerous applications, but the demand for ever more performance remains unabated. Scaling CNN computations to larger clusters is generally done by distributing tasks in batch mode using methods such as distributed synchronous SGD. Among the issues with this approach is that, to make the distributed cluster work with high utilization, the workload distributed to each node must be large; this implies nontrivial growth in the SGD mini-batch size. In this article we propose a framework, called FPDeep, which uses a hybrid of model and layer parallelism to configure distributed reconfigurable clusters to train CNNs. This approach has numerous benefits. First, the design does not suffer from performance loss due to batch size growth. Second, work and storage are balanced among nodes through novel workload and weight partitioning schemes. Part of the mechanism is the surprising finding that it is preferable to store excess weights in neighboring devices rather than in local off-chip memory. Third, the entire system is a fine-grained pipeline. This leads to high parallelism and utilization and also minimizes the time that features need to be cached while waiting for back-propagation. As a result, storage demand is reduced to the point where only on-chip memory is used for the convolution layers. And fourth, we find that the simplest topology, a 1D array, is preferred for interconnecting the FPGAs thus enabling widespread applicability. We evaluate FPDeep with the Alexnet, VGG-16, and VGG-19 benchmarks. Results show that FPDeep has good scalability to a large number of FPGAs, with the limiting factor being the FPGA-to-FPGA bandwidth. But with 250 Gb/s bidirectional bandwidth per FPGA, which is easily supported by current generation FPGAs, FPDeep performance shows linearity up to 100 FPGAs. Energy efficiency is evaluated with respect to GOPs/J. FPDeep provides, on average, 6.4× higher energy efficiency than comparable GPU servers. Tong Geng, Ang Li 0006, Xi Jin 0002, Martin C. Herbordt |
IEEE Trans. Computers | 4 |
| 2019 | Accelerating AP3M-Based Computational Astrophysics Simulations with Reconfigurable ClustersabstractIn this paper, we present a case study of using a reconfigurable computing cluster to accelerate AP3M-based computational astrophysics simulations. AP3M is an adaptive particle-particle, particle-mesh method. Many computational astrophysics simulations are based on this method. AP3M can dynamically and adaptively apply computational resources non-uniformly to emphasize regions of interest. Therefore, AP3M can be faster and more energy-efficient than the traditional P3M (particle-particle, particle-mesh) approach. However, the dynamic and pointer-based data structure used by AP3M makes it extremely difficult to accelerate with FPGAs. In this work, we use a custom data structure and hardware kernel to overcome these challenges. All CPU-based dynamic and pointer-based tasks are mapped to FPGAs. Our experiments show that a single FPGA outperforms a Xeon E5-2660 CPU server (8 cores) by from 21x to 23x depending on problem size and data distribution. Tong Geng, Xi Jin 0002, Martin C. Herbordt |
ASAP | 3 |
| 2019 | FP-AMR: A Reconfigurable Fabric Framework for Adaptive Mesh Refinement ApplicationsabstractAdaptive mesh refinement (AMR) is one of the most widely used methods in High Performance Computing accounting a large fraction of all supercomputing cycles. AMR operates by dynamically and adaptively applying computational resources non-uniformly to emphasize regions of the model as a function of their complexity. Because AMR generally uses dynamic and pointer-based data structures, acceleration is challenging, especially in hardware. As far as we are aware there has been no previous work published on accelerating AMR with FPGAs. In this paper, we introduce a reconfigurable fabric framework called FP-AMR. The work is in two parts. In the first FP-AMR offloads the bulk per-timestep computations to the FPGA; analogous systems have previously done this with GPUs. In the second part we show that the rest of the CPU-based tasks-including particle mesh mapping, mesh refinement, and coarsening-can also be mapped efficiently to the FPGA. We have evaluated FP-AMR using the widely used program AMReX and found that a single FPGA outperforms a Xeon E5-2660 CPU server (8 cores) by from 21x -23x depending on problem size and data distribution. Tong Geng, Xi Jin 0002, Martin C. Herbordt |
FCCM | 3 |
| 2019 | CINT - An Energy-efficient Mixed-signal In-Memory CNN Accelerator Based on NOR Flash MemoryabstractConvolutional neural network (CNN) is a power-hungry and resource-consuming application, which makes it hard to deploy on end devices. We propose a method to perform convolution operations in NOR flash memory. Experiment results show that our method has great performance and high energy efficiency. Linfeng Tao, Teng Tian, Zikun Xiang, Xi Jin 0002, Zhengda Li, Chenxia Li |
MobiSys | 6 |
| 2018 | A Real-Time Learning-Based Super-Resolution System Using Direct Simple FunctionsabstractThis paper proposes a real-time super-resolution (SR) system. The proposed system performs a fast SR algorithm that generates a high-resolution image from a low-resolution image using direct regression functions. The system implemented on a Xilinx Virtex 7 field programmable gate array achieves output resolution of 3840 × 2160 (UHD) at 200 fps and 2000Mpixels/s throughput. Experimental results show that the proposed system provides high image quality for real-time applications. Daolu Zha, Xi Jin 0002 |
ASAP | 2 |
| 2018 | An efficient resource-optimized learning prefetcher for solid state drivesabstractIn recent years, solid-state drives (SSDs) have been widely deployed in modern storage systems. To increase the performance of SSDs, prefetchers for SSDs have been designed both at operating system (OS) layer and flash translation layer (FTL). Prefetchers in FTL have many advantages like OS-independence, easy-using, and compatibility. However, due to the limitation of computing capabilities and memory resources, existing prefetchers in FTL merely employ simple sequential prefetching which may incur high penalty cost for I/O access stream with complex patterns. In this paper, an efficient learning prefetcher implemented in FTL is proposed. Considering the resource limitation of SSDs, a learning algorithm based on Markov chains is employed and optimized so that high hit ratio and low penalty cost can be achieved even for complex access patterns. To validate our design, a simulator with the prefetcher is designed and implemented based on Flashsim. The TPC-H benchmark and an application launch trace are tested on the simulator. According to experimental results of the TPC-H benchmark, more than 90% of memory cost can be saved in comparison with a previous design at OS layer. The hit ratio can be increased by 24.1% and the number of times of misprefetching can be reduced by 95.8% in comparison with the simple sequential prefetching strategy. Xi Jin 0002, Linfeng Tao, Shuaizhi Guo, Zikun Xiang, Teng Tian |
DATE | 2 |
| 2016 | RP-Ring: A Heterogeneous Multi-FPGA Accelerating Solution for N-Body SimulationsabstractWe propose an heterogeneous multi-FPGA accelerating solution, which is called as RP-ring (Reconfigurable Processor ring), for direct-summation N-body simulation. In this solution, we try to use existing FPGA boards rather than design new specialized boards to reduce cost. It can be expanded conveniently with any available FPGA board and only requires quite low communication bandwidth between FPGA boards. The communication protocol is simple and can be implemented with limited hardware/software resource. In order to prevent the slowest board from dragging the overall performance down, we build a mathematical model to decompose workload among FPGAs. The model divide workload based on the logic resource, memory access bandwidth and communication bandwidth of each FPGA chip. We apply the solution in astrodynamics simulation and achieve two orders of magnitude speedup compared with CPU implementations. Xi Jin 0002, Chuanjun Wang, Linlin Zheng |
FCCM | 2 |
| 2016 | An FPGA-SOC Based Accelerating Solution for N-body Simulations in MOND (Abstract Only)abstractModified Newtonian dynamics (MOND) has shown a great success as a modified-potential theory of gravity. In this paper, we present a highly integrated accelerating solution for N-body MOND simulations. By using the FPGA-SoC, which integrates both FPGA and SOC (system on chip) in one chip, our solution exhibits potential for better performance, higher integration, and lower power consumption. To handle the calculation bottleneck of potential summation, on one hand, we develop a strategy to simplify the pipeline, in which the square calculation task is conducted by the DSP48E1 of Xilinx 7 series FPGAs, so as to reduce the logic resource consumption of each pipeline; on the other hand, advantages of particle-mesh scheme are taken to overcome the bottleneck on bandwidth. Our experiment results show that 2 more pipelines can be integrated in Zynq-7020 FPGA-SoC with the simplified pipeline, and the bandwidth requirement is reduced significantly. Furthermore, our accelerating solution has a full range of advantages over different processors. Compared with GPU, our work is about better in both performance per Watt and performance per cost. Xi Jin 0002, Chuanjun Wang |
FPGA | 3 |
| 2016 | an Extensible Heterogeneous Multi-FPGA Framework for Accelerating N-body Simulation (Abstract Only)abstractN-body simulation plays a significant role in scientific research and engineering development. Direct-summation N-body algorithms compute the particle interaction in an exact way, but this algorithm have a computational complexity of $O(N^2)$. To simulate a large system efficiently and flexibly, lots of high performance implementations on FPGA have been developed. Xi Jin 0002 |
FPGA | 3 |
| 2016 | An Improved Global Stereo-Matching on FPGA for Real-Time Applications (Abstract Only)abstractA real-time global stereo matching algorithm is implemented on FPGA. Stereo matching is frequently used in stereo vision systems, e.g. for stereo vision applications like objects detection and autonomous vehicles. Global algorithms perform much more significant than local algorithms, but global algorithms are not implemented on FPGA by reason of rely on the high-end hardware resources. In this implementation the stereo pairs are divided into blocks, the hardware resources are reduced by processing one block once. The hardware implementation is based on a Xilinx®Kintex 7 FPGA. Experiment results show the implementation performances significant and 30 [email protected] is achieved. Daolu Zha, Xi Jin 0002 |
FPGA | 2 |
| 2016 | FPGA acceleration of TreePM N-body simulations for Modified Newton DynamicsabstractIn the field of high-performance and energy-efficient scientific computing, FPGAs are promising candidates. In this paper, we show a case study of applying FPGA-based system for Modified Newtonian Dynamics (MOND) cosmological simulation. The numerical simulation is based on TreePM algorithm of N-body problem. TreePM algorithm combines the PM (Particle Mesh) method on large scales with a tree method to provide fine resolution. So two kinds of accelerate module are integrated in FPGA. We leverage the customizability of the FPGA on-chip memory to construct a dynamic reconfigurable cache in order to enhance the performance of the interface between accelerate modules and external memory. Our solution achieve 37.3 times performance than CPU implementation. Linlin Zheng, Xi Jin 0002, Chuanjun Wang |
FPT | 3 |
| 2016 | Fixed-ratio DXT format Frame Buffer Compressor for mobile graphics systemsabstractMost of Frame Buffer Compression (FBC) methods are based on variable-length compression algorithm. The proposed fixed-ratio frame buffer compressor, which is based on S3 texture compression (S3TC) algorithm, compresses the recently rendered image into DXT format so that it can be employed as input compressed texture. We have optimized the frame buffer compressor to achieve real time performance and FPGA implementation, which occupies 406 slice registers and 1133 slice LUT in Xilinx Spartan6 FPGA. The compressor consumes 10% additional silicon space, but it reduces the bandwidth consumption to 16.8%-45.8% compared to that of 50-90% by the variable-length compression methods. Yuzhi Zhou, Xi Jin 0002 |
FPT | 2 |
| 2015 | Design of a Distributed Compressor for Astronomy SSDabstractSSD (solid state device) has shown a great potential in astronomy data storage. Data compression is an essential task to obtain higher storage density and bandwidth. This paper proposes a distributed compressor customized for FPGA-based astronomy SSD. Our data-driven compressor cope with astronomy data in the unit of byte, two compression algorithms, run length and length-limited huffman are utilized, a distributed length-limited huffman encoder for SSD is further developed to reduce the latency. Experimental results indicate that our proposed compressor achieves a 1GB/s bandwidth with less than 2500 LUTs utilized while the compression ratio is only 10% lower than Gzip level9. Xi Jin 0002, Xueliang Du |
FCCM | 2 |
| 2014 | A Multi-phase Clock Time-to-Digital Convertor Based on ISERDES ArchitectureabstractThe time-to-digital converter(TDC) aims to mark an accurate timestamp at the time of input signal comes. The Multi-phase Clock sampling method is an usual way to map the TDC into an FPGA. Traditionally, this method provides a medium accuracy and low resources occupation. In this paper, we present a new architecture of TDC base on the 2-ISERDES in the SelectIO, rather than utilizing the Slice resources by the old way. The ISERDESes based TDC is equivalent to a 8 equidistant phase-shifted clocks TDC, with maximum clock frequency 900MHz. The least significant bit(LSB) is 139ps, which is 445% better than traditional architecture. Xi Jin 0002, Shaoping Chu, Xue Ben |
FCCM | 3 |