EDBT 2026 Demo / reviewers in the wild / expert
Teng Tian
dblp:218/1159
· DBLP profile ↗
10ranked-venue papers
3as first author
7since 2021 · last 2024
0000-0002-0594-5957ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorComputer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | RingTK: A Ring, Parallel and High Performance Top-K Sorter on FPGAabstractGetting the K largest/smallest elements from$N$inputs is one of the essential operations in many applications. In this article, we propose RingTK, a ring, parallel, and high performance Top-K sorter implemented on FPGA. We use a priority queue as the basic processing unit and design a Top-K sorter based on a ring topology with a global maximum module(GMM) to obtain good scalability and high performance. Based on the ring topology, we design Ring Multiplexers (RMUX) and modify the GMM to enable RingTK to efficiently handle different K-sizes and concurrent tasks. We design a encoder to flexibly deal with different data formats and max/min Top-K tasks. Finally, we implement the proposed architecture on the Xilinx XCVU37P FPGA. The results show that the proposed architecture has good scalability, and the throughput and parallelism are approximately linear. We can achieve a throughput of 35.38GB/s with 32 PQs that exceed existing literature with K=8160. Huawen Liang, Qizhe Wu, Wei Yuan 0006, Teng Tian, Xi Jin 0002 |
FCCM | 4 |
| 2024 | Artificial Intelligence-Guided Fully-Automatic Renal Segmentation
Teng Tian, Yidong Gu, Ruwang Jiao, Tao Peng 0013 |
PRICAI (3) | 2 |
| 2022 | FP-GNN: Adaptive FPGA accelerator for Graph Neural Networks
Teng Tian, Letian Zhao, Qizhe Wu, Wei Yuan 0006, Xi Jin 0002 |
Future Gener. Comput. Syst. | 1 |
| 2022 | G-NMP: Accelerating Graph Neural Networks with DIMM-based Near-Memory Processing
Teng Tian, Letian Zhao, Xuecang Zhang, Fangmin Lu, Xi Jin 0002 |
J. Syst. Archit. | 1 |
| 2022 | QEGCN: An FPGA-based accelerator for quantized GCNs with edge-level parallelism
Wei Yuan 0006, Teng Tian, Qizhe Wu, Xi Jin 0002 |
J. Syst. Archit. | 2 |
| 2021 | Malicious Domain Detection on Imbalanced Data with Deep Reinforcement Learning
Fangfang Yuan, Teng Tian, Yanmin Shang, Yuhai Lu, Yanbing Liu 0007, Jianlong Tan |
ICONIP (4) | 2 |
| 2021 | A Gather Accelerator for GNNs on FPGA PlatformabstractGraph Neural Networks (GNNs) have emerged as the state-of-the-art deep learning model for representation learning on graphs. GNNs mainly include two phases with different execution patterns. The Gather phase, depends on the structure of the graph, presenting a sparse and irregular execution pattern. The Apply phase, acts like other neural networks, showing a dense and regular execution pattern. It is challenging to accelerate GNNs, due to irregular data communication to gather information within the graph. To address this challenge, hardware acceleration for Gather phase is critical. The purpose of this research is to design and implement an FPGA-based accelerator for Gather phase. It achieves excellent performance on acceleration and energy efficiency. Evaluation is performed using a Xilinx VCU128 FPGA with three commonly-used datasets. Compared to the state-of-the-art software framework running on Intel Xeon CPU and NVIDIA P100 GPU, our work achieves on average 101.28× speedup with 75.27× dynamic energy reduction and average 12.27× speedup with 45.56× dynamic energy reduction, respectively. Wei Yuan 0006, Teng Tian, Huawen Liang, Xi Jin 0002 |
ICPADS | 2 |
| 2020 | Exploration of Memory Access Optimization for FPGA-based 3D CNN AcceleratorabstractThree-dimensional convolutional networks (3D CNNs) are used efficiently in various video recognition applications. Compared to traditional 2D CNNs, extra temporal dimension causes 3D CNNs more computationally intensive and to have a larger memory footprint. Therefore, the memory optimization is extremely crucial in this case. This paper presents a design space exploration of memory access optimization for FPGA-based 3D CNN accelerator. We present a non-overlapping data tiling method for contiguous off-chip memory access and explore on-chip data reuse opportunity by leveraging different loop ordering strategies. We propose a hardware architecture design which can flexibly support different loop ordering strategies for each 3D CNN layer. With the help of hardware/software co-design, we can provide the optimal configuration toward an energy-efficient and high-performance accelerator design. According to the experiments on AlexNet, VGG16, and C3D, our optimal model reduces up to 84% DRAM accesses and 55% energy consumption on C3D compared to a baseline model, and demonstrates state-of-the-art performance compared to prior FPGA implementations. Teng Tian, Xi Jin 0002, Letian Zhao |
DATE | 1 |
| 2019 | CINT - An Energy-efficient Mixed-signal In-Memory CNN Accelerator Based on NOR Flash MemoryabstractConvolutional neural network (CNN) is a power-hungry and resource-consuming application, which makes it hard to deploy on end devices. We propose a method to perform convolution operations in NOR flash memory. Experiment results show that our method has great performance and high energy efficiency. Linfeng Tao, Teng Tian, Zikun Xiang, Xi Jin 0002, Zhengda Li, Chenxia Li |
MobiSys | 3 |
| 2018 | An efficient resource-optimized learning prefetcher for solid state drivesabstractIn recent years, solid-state drives (SSDs) have been widely deployed in modern storage systems. To increase the performance of SSDs, prefetchers for SSDs have been designed both at operating system (OS) layer and flash translation layer (FTL). Prefetchers in FTL have many advantages like OS-independence, easy-using, and compatibility. However, due to the limitation of computing capabilities and memory resources, existing prefetchers in FTL merely employ simple sequential prefetching which may incur high penalty cost for I/O access stream with complex patterns. In this paper, an efficient learning prefetcher implemented in FTL is proposed. Considering the resource limitation of SSDs, a learning algorithm based on Markov chains is employed and optimized so that high hit ratio and low penalty cost can be achieved even for complex access patterns. To validate our design, a simulator with the prefetcher is designed and implemented based on Flashsim. The TPC-H benchmark and an application launch trace are tested on the simulator. According to experimental results of the TPC-H benchmark, more than 90% of memory cost can be saved in comparison with a previous design at OS layer. The hit ratio can be increased by 24.1% and the number of times of misprefetching can be reduced by 95.8% in comparison with the simple sequential prefetching strategy. Xi Jin 0002, Linfeng Tao, Shuaizhi Guo, Zikun Xiang, Teng Tian |
DATE | 6 |