Xingyu Tian

dblp:268/9454 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 4 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 An Efficient and Scalable Hardware Architecture for Number Theoretic Transform on FPGA with Design Automation
abstract
Fully Homomorphic Encryption (FHE) has become a promising approach to protecting data privacy in emerging application scenarios. Unfortunately, FHE suffers from significant processing speed degradation compared to plaintext computation, with one of the primary bottlenecks being the time-consuming Number Theoretic Transform (NTT). Therefore, accelerating NTT to accommodate various FHE parameters is crucial to advancing FHE towards practical use. With highly reconfigurable and performant logical fabrics, Field Programmable Gate Arrays (FPGAs) have exhibited great potential in NTT acceleration. By decomposing large-point NTT with strong data dependency into independent and simple small-point NTTs, the emerging Ten-step NTT (TNTT) algorithms intuitively enable higher parallelism and thereby have the potential to explore better performance compared to traditional algorithms. However, our quantitative analysis reveals that TNTT exhibits significant performance degradation as parallelism increases due to additional varying-size transpositions and Hadamard products. This paper proposes AutoNest, an efficient and scalable hardware architecture, along with an accelerator auto-generation framework for TNTT. The proposed hardware architecture maximizes performance by 1) adopting a 2D block decomposition dataflow to address critical path delays in transpose logic, thereby improving clock frequency. 2) integrating algorithm-level costfree twiddle factor fusion to reduce the number of modular multiplications in Hadamard products, thereby allowing higher parallelism on chip. Moreover, we also deliver an accelerator generation framework conducting automated design space exploration to elaborate a performant TNTT architecture under the target FPGAs' resource budget for user-defined FHE parameters. Experimental results on the AMD-Xilinx U280 FPGA demonstrate that NTT accelerators generated by AutoNest achieve an average speedup of$2.31 \times$compared to prior designs.
Yilan Zhu, Geng Yang 0001, Xingyu Tian, Dilshan Kumarathunga, Liang Kong 0005, Xianglong Deng, Shengyu Fan, Guang Fan 0001, Guiming Shi, Bo Zhang 0098, Yisong Chang, Shoumeng Yan, Zhenman Fang, Mingzhe Zhang 0005
HPCA3
2025 MAD-HiSpMV: Matrix Adaptive Design with Hybrid Row Distribution for Imbalanced SpMV Acceleration on FPGAs
abstract
Sparse Matrix–Vector Multiplication (SpMV) is fundamental in numerous applications such as scientific computing, Machine Learning (ML), and graph analytics. While recent studies have made tremendous progress in accelerating SpMV on HBM-equipped FPGAs, there are still multiple remaining challenges to accelerate imbalanced SpMV where the distribution of nonzeros in the sparse matrix is imbalanced across different rows. These include (1) imbalanced workload distribution among the parallel Processing Elements (PEs), (2) long-distance dependency for floating-point accumulation on the output vector, (3) a new bottleneck due to the often-overlooked dense vectors’ off-chip access after the SpMV acceleration, and (4) sub-optimal performance of generic accelerators for various types of sparse matrices. (5) Additionally, ML workloads often consist of both SpMV and General Matrix–Vector Multiplication (GeMV), which suffer from kernel switching inefficiencies. To address those challenges, we propose MAD-HiSpMV to accelerate imbalanced SpMV on HBM-equipped FPGAs with the following novel solutions: (1) a hybrid row distribution network to enable both inter-row and intra-row distribution for better balance, (2) a fully pipelined floating-point accumulation on the output vector using a combination of an adder chain and register-based circular buffer, (3) matrix adaptive design configurations generated by our automation framework via Design Space Exploration (DSE) to maximize performance for the given matrix, and (4) a GeMV overlay built into the same kernel for efficient acceleration of mixed workloads. Experimental results demonstrate that the DSE-picked configuration of MAD-HiSpMV achieves a geomean speedup of 1.3× (up to 2.12×) for the SpMV benchmark matrices and achieves a geomean 1.15× (up to 1.54×) better performance per watt, when compared to state-of-the-art generic designs. For the SpMV benchmark matrices, compared to Intel MKL running on a 24-core Xeon Silver 4214 CPU, MAD-HiSpMV achieves a geomean speedup of 8.80×. Compared to cuSparse running on an Nvidia GTX 1080ti GPU, MAD-HiSpMV achieves a geomean of 2.57× better performance per watt. Additionally, a GeMV overlay built into MAD-HiSpMV achieves a peak throughput of 156.7 GFLOPS, which is 2.64× better than the Vitis L2 GeMV benchmark on U280, and performs 2.7× better for an end-to-end mixed workload, when compared to Intel MKL running on a 24-core Xeon Silver 4214 CPU. MAD-HiSpMV is available at https://github.com/SFU-HiAccel/HiSpMV .
Manoj B. Rajashekar, Akhil Raj Baranwal, Xingyu Tian, Zhenman Fang
ACM Trans. Reconfigurable Technol. Syst.3
2024 HiSpMV: Hybrid Row Distribution and Vector Buffering for Imbalanced SpMV Acceleration on FPGAs
abstract
Sparse matrix-vector multiplication (SpMV) is a fundamental operation in numerous applications such as scientific computing, machine learning, and graph analytics. While recent studies have made great progress in accelerating SpMV on HBM-equipped FPGAs, there are still multiple remaining challenges to efficiently accelerate imbalanced SpMV where the distribution of non-zeros in the sparse matrix is imbalanced across different rows. First, the imbalanced workload distribution among the parallel processing elements (PEs) leads to PE under-utilization and performance degradation. Second, the read-after-write dependency of the long-latency floating-point accumulation on the output vector causes pipeline stalls inside the PE, and existing scheduling solutions for balanced matrices no longer work effectively for imbalanced ones. Third, the memory access latency for the often overlooked input vector becomes a new performance bottleneck after the SpMV acceleration.
Manoj B. Rajashekar, Xingyu Tian, Zhenman Fang
FPGA2
2024 PASTA: Programming and Automation Support for Scalable Task-Parallel HLS Programs on Modern Multi-Die FPGAs
abstract
In recent years, the adoption of FPGAs in datacenters has increased, with a growing number of users choosing High-Level Synthesis (HLS) as their preferred programming method. While HLS simplifies FPGA programming, one notable challenge arises when scaling up designs for modern datacenter FPGAs that comprise multiple dies. The extra delays introduced due to die crossings and routing congestion can significantly degrade the frequency of large designs on these FPGA boards. Due to the gap between HLS design and physical design, it is challenging for HLS programmers to analyze and identify the root causes, and fix their HLS design to achieve better timing closure. Recent efforts have aimed to address these issues by employing coarse-grained floorplanning and pipelining strategies on task-parallel HLS designs where multiple tasks run concurrently and communicate through FIFO stream channels. However, many applications are not streaming friendly and many existing accelerator designs heavily rely on buffer channel based communication between tasks. In this work, we take a step further to support a task-parallel programming model where tasks can communicate via both FIFO stream channels and buffer channels. To achieve this goal, we design and implement the PASTA framework, which takes a large task-parallel HLS design as input and automatically generates a high-frequency FPGA accelerator via HLS and physical design co-optimization. Our framework introduces a latency-insensitive buffer channel design, which supports memory partitioning and ping-pong buffering while remaining compatible with vendor HLS tools. On the frontend, we provide an easy-to-use programming model for utilizing the proposed buffer channel; while on the backend, we implement efficient placement and pipelining strategies for the proposed buffer channel. To validate the effectiveness of our framework, we test it on four widely used Rodinia HLS benchmarks and two real-world accelerator designs and show an average frequency improvement of 25%, with peak improvements of up to 89% on AMD/Xilinx Alveo U280 boards compared to Vitis HLS baselines.
Moazin Khatti, Xingyu Tian, Ahmad Sedigh Baroughi, Akhil Raj Baranwal, Yuze Chi, Licheng Guo, Jason Cong, Zhenman Fang
ACM Trans. Reconfigurable Technol. Syst.2
2023 PASTA: Programming and Automation Support for Scalable Task-Parallel HLS Programs on Modern Multi-Die FPGAs
abstract
In recent years, there has been increasing adoption of FPGAs in datacenters as hardware accelerators, where a large population of end users are software developers. While high-level synthesis (HLS) facilitates software programming, it is still challenging to scale large accelerator designs on modern datacenter FPGAs that often consist of multiple dies and memory banks. More specifically, routing congestion and extra delays on these multi-die FPGAs often cause timing closure issues and severe frequency degradation at the physical design level, which are difficult to digest and optimize for high-level programmers using HLS. One promising approach to mitigate such issues is to develop a high-level task-parallel programming model with HLS and physical design co-optimization. Unfortunately, existing studies only support a programming model where tasks communicate with each other via FIFOs, while many applications are not streaming friendly and many existing accelerator designs heavily rely on buffer based communication between tasks. In this paper, we take a step further to support a task-parallel programming model where tasks can communicate via both FIFOs and buffers. To achieve this goal, we design and implement the PASTA framework, which takes a large task-parallel HLS design as input and automatically generates a high-frequency FPGA accelerator via HLS and physical design co-optimization. First, we design a decoupled latency-insensitive buffer channel that supports memory partitioning and ping-pong buffering, which is compatible with the vendor Vitis HLS compiler. In the frontend, we develop an easy-to-use programming interface to allow end users to use our buffer channel in their applications. In the backend, we provide automatic coarse-grained floorplanning and pipelining for designs that use our proposed buffer channel. We test PASTA on a set of task-parallel HLS designs that use buffers for task communication and show an average of 36% (up to 54%) frequency improvement for large design configurations.
Moazin Khatti, Xingyu Tian, Yuze Chi, Licheng Guo, Jason Cong, Zhenman Fang
FCCM2
2023 Visible-Thermal Person Reidentification in Visual Internet of Things With Random Gray Data Augmentation and a New Pooling Mechanism
abstract
Visible–thermal person reidentification (VT-ReID) is an emerging cross-modality matching problem, which aims to identify the same person across the daytime visible modality and nighttime thermal modality in the Internet of Things. Existing cutting-edge approaches consistently attempt to exploit image generation technique to generate cross-modality images or design various feature-level constraints to align feature distribution of heterogeneous data. However, color variations originating from the different imaging processes of spectrum cameras remain unsolved, which leads to suboptimal feature representations. In this article, we present a simple but very effective data augmentation method named Random Gray for the cross-modality matching task. Given a training sample, Random Gray randomly selects a rectangular region and translates it to grayscale. In this process, training images with fusing various levels of visible and grayscale information are generated, thereby reducing the risk of overfitting and making the model robust to color variations. Besides, we introduce a novel pooling method called softpooling to retain more information in the reduced activation maps. With softpooling layer, the network can learn more discriminative person features and further boost its retrieval performance. We conduct extensive experiments on publicly available cross-modality Re-ID data sets (SYSU-MM01 and RegDB) to demonstrate the effectiveness of our proposed method. Experimental results show that Random Gray and softpooling strategies yield significant accuracy improvement, and they can be utilized as training tricks for further VT-ReID research.
Xingyu Tian, Houfu Peng, Daoxun Xia
IEEE Internet Things J.3
2023 A Ground-Roll Separation Method Based on Neural Networks With Morphological Similarity Loss
abstract
Ground-roll is a typical Rayleigh-type interference noise in field seismic data, which is characterized by low frequency, low velocity and high amplitude. Since it will interfere effective seismic signals and severely degrade the signal-to-noise ratio of observed seismic records, many approaches have been developed for ground-roll attenuation or separation. In this letter, we proposed an improved ground-roll separation algorithm through the combination of deep learning based low-frequency generation and dictionary learning based low-frequency reconstruction. Moreover, to utilize the inter-band morphological similarity prior in seismic response, we introduce the morphological similarity constraint into the learning approach of pseudo low-frequency generation networks. Experiments demonstrate that compared to previous methods, the introduced the morphological similarity loss can effectively improve the quality of generated pseudo low-frequency signals, which results in better low-frequency reflection reconstruction and ground-roll separation performances.
Xingyu Tian, Yile Ao, Yanda Li, Wenkai Lu
IEEE Geosci. Remote. Sens. Lett.1
2023 Improved Seismic Residual Diffracted Multiple Suppression Method Based on Object Detection and Image Segmentation
abstract
Seismic multiple is one of the most common noises in marine seismic data, which heavily affects subsequent processing and interpretation. To eliminate the influence of seismic multiples, many methods have been developed, while surface-related multiple elimination (SRME) is one of the most widely deployed methods. However, results of SRME always contain a few strong residual diffracted multiples (RDMs) in practice because of the unprecise prediction of diffracted multiples compared to reflection multiples. If we try to apply further multiple suppression methods to SRME results, it not only tends to damage the signals, but also spends lots of unnecessary computations where there is no RDM. In this article, we propose an improved RDM suppression method based on object detection and image segmentation. First, we employ an object detection network to locate bounding boxes containing RDMs in the SRME results. Then a threshold-based image segmentation method is utilized to identify regions of strong RDMs in the detected boxes. According to the segmentation results, parameters for weak multiples and strong multiples are provided for the adaptive multiple subtraction (AMS) in different regions to generate different results. At last, we combine the suppression results of strong RDMs and weak RDMs as the final results. Application on field data demonstrates that our method is able to suppress RDMs with little loss of signal.
Xingyu Tian, Wenkai Lu, Yanda Li, Mingrui Zhong, Hongxun Pan, Bowu Jiang
IEEE Trans. Geosci. Remote. Sens.1
2023 TAPA: A Scalable Task-parallel Dataflow Programming Framework for Modern FPGAs with Co-optimization of HLS and Physical Design
abstract
In this article, we propose TAPA, an end-to-end framework that compiles a C++ task-parallel dataflow program into a high-frequency FPGA accelerator. Compared to existing solutions, TAPA has two major advantages. First, TAPA provides a set of convenient APIs that allows users to easily express flexible and complex inter-task communication structures. Second, TAPA adopts a coarse-grained floorplanning step during HLS compilation for accurate pipelining of potential critical paths. In addition, TAPA implements several optimization techniques specifically tailored for modern HBM-based FPGAs. In our experiments with a total of 43 designs, we improve the average frequency from 147 MHz to 297 MHz (a 102% improvement) with no loss of throughput and a negligible change in resource utilization. Notably, in 16 experiments, we make the originally unroutable designs achieve 274 MHz, on average. The framework is available at https://github.com/UCLA-VAST/tapa and the core floorplan module is available at https://github.com/UCLA-VAST/AutoBridge
Licheng Guo, Yuze Chi, Jason Lau, Linghao Song, Xingyu Tian, Moazin Khatti, Weikang Qiao, Jie Wang 0022, Ecenur Ustun, Zhenman Fang, Zhiru Zhang, Jason Cong
ACM Trans. Reconfigurable Technol. Syst.5
2023 SASA: A Scalable and Automatic Stencil Acceleration Framework for Optimized Hybrid Spatial and Temporal Parallelism on HBM-based FPGAs
abstract
Stencil computation is one of the fundamental computing patterns in many application domains such as scientific computing and image processing. While there are promising studies that accelerate stencils on FPGAs, there lacks an automated acceleration framework to systematically explore both spatial and temporal parallelisms for iterative stencils that could be either computation-bound or memory-bound. In this article, we present SASA, a scalable and automatic stencil acceleration framework on modern HBM-based FPGAs. SASA takes the high-level stencil DSL and FPGA platform as inputs, automatically exploits the best spatial and temporal parallelism configuration based on our accurate analytical model, and generates the optimized FPGA design with the best parallelism configuration in TAPA high-level synthesis C++ as well as its corresponding host code. Compared to state-of-the-art automatic stencil acceleration framework SODA that only exploits temporal parallelism, SASA achieves an average speedup of 3.41× and up to 15.73× speedup on the HBM-based Xilinx Alveo U280 FPGA board for a wide range of stencil kernels.
Xingyu Tian, Zhifan Ye, Alec Lu, Licheng Guo, Yuze Chi, Zhenman Fang
ACM Trans. Reconfigurable Technol. Syst.1
2022 Real-time sentiment analysis of students based on mini-Xception architecture for wisdom classroom
abstract
Abstract Sentiment analysis has a wide application prospect in business, medicine, security and other fields, which provides a new perspective for the development of education. Students' sentiment data play an important role in the evaluation of teachers' teaching quality and students' learning effect, and provide a basis for the implementation of effective learning intervention. However, most of the research is to obtain the real‐time learning status of students in the classroom through teachers' naked eye observation and students' text feedback, which will lead to some problems such as incomplete feedback content and delayed feedback analysis. Based on the mini‐Xception framework, this article implements the real‐time identification and analysis of student sentiment in classroom teaching, and the degree of student engagement is analyzed according to the teaching events triggered by teacher to provide reasonable suggestions for subsequent teaching progress. The experimental results show that the mini‐Xception model trained by FER2013 data sets has high recognition accuracy for the real‐time detection of seven student sentiments, and the average accuracy is 76.71%. Compared with text feedback, it can assist teachers in understanding student learning states in time so that they can take corresponding actions, and realize the real‐time performance of wisdom classroom teaching information feedback, the high efficiency of information transmission, and the intelligence of information processing.
Xingyu Tian, Shengnan Tang, Daoxun Xia
Concurr. Comput. Pract. Exp.1
2022 Data Cleansing for Salt Dome Dataset With Noise Robust Network on Segmentation Task
abstract
Noisy labels seriously degrade the performance of the deep learning models. Especially on segmentation tasks, labels are represented at pixel level, and therefore, it is easier to generate them as noisy labels. In this letter, we aim to cleanse the dataset which contains noisy labels using Kullback–Leibler (KL) divergence algorithm and noise robust loss function. Here, we regard the whole dataset as noisy labels and separate noisy labels into two different parts. One is the strong noisy labels where there are the most noisy labels, which is referred to as incorrect labels, and the other is The weak noisy labels where pixels are incorrect in some parts of the image. The KL algorithm is more efficient in removing the labels which contain the most noise. At the same time, we use noise robust model to remove the labels where pixels are inaccurate only in small parts of the image. Noise robust loss function is robust to noisy labels, and thus, we consider noisy labels when intersection over union (IoU) is lower than a constant threshold despite using noise robust loss function. The effectiveness of our proposed method is assessed using Kaggle’s TGS Salt Identification Challenge dataset. We have demonstrated that using our method, segmentation performance is increased without heavy noisy labels.
Youdam Chung, Wenkai Lu, Xingyu Tian
IEEE Geosci. Remote. Sens. Lett.3
2022 Improved Anomalous Amplitude Attenuation Method Based on Deep Neural Networks
abstract
In seismic exploration, seismic data usually contain anomalous amplitude noise whose high energy may affect the results of subsequent processing steps. In industry, this kind of noise is generally suppressed using the anomalous amplitude attenuation (AAA) method. The AAA method essentially suppresses abnormal amplitude noise using a median filter in the time–frequency domain. This makes its performance heavily dependent on the parameters, especially the window width of the median filter. Thus, we propose an improved anomalous amplitude attenuation (IAAA) method based on deep neural networks. The IAAA method contains two steps. In the first step, deep neural networks are used to detect the locations and the widths of noise regions. In the second step, the noise information (locations and widths) obtained at the previous step is exploited to apply the AAA method with more appropriate parameters to each noisy region. Compared with the conventional AAA method, the IAAA method can suppress the noise more effectively and preserve signals better. The experiments on both synthetic data and field data demonstrate that our method outperforms the conventional AAA method.
Xingyu Tian, Wenkai Lu, Yanda Li
IEEE Trans. Geosci. Remote. Sens.1