Xingyun Qi

dblp:88/8790 · DBLP profile ↗
← Back
17ranked-venue papers
2as first author
15since 2021 · last 2026
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 15 · 14 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 TL-Sort: A Fully Pipelined Hardware Architecture for Sorting Without Run-Drain Stalls
Hai Cao, Puguang Liu, Zhang Luo, Jihang Wang, Xingyun Qi
APPT6
2026 A 7×25Gb/s Transceiver Using Codebook-Based CNRZ-7 for High-Density Transmission
Ruixiao Kuai, Fangxu Lv, Xingyun Qi, Liangyong Yuan, Lizhou Wu, Bohui Bai, Ruotian Yin
ISCAS5
2026 A Capacitor-Less Current-Feedback LDO with Mismatch Cancellation Using DEM and Chopping for Distributed Power Management
Chengzhuo Zhao, Fangxu Lv, Xingyun Qi, Kewei Xin, Wenchen Wang, Jiliang Liu
ISCAS5
2026 Self -adaptive and topology-aware broadcast leveraging collective offload on Tianhe express interconnect
Chongshan Liang, Xingyun Qi, Dongsheng Li 0001
J. Parallel Distributed Comput.3
2026 CXL-DMSim: A Full-System CXL Disaggregated Memory Simulator With Comprehensive Silicon Validation
abstract
Compute eXpress Link (CXL) has emerged as a key enabler of memory disaggregation for future heterogeneous computing systems to expand memory on-demand and improve resource utilization. However, CXL is still in its infancy stage and lacks commodity products on the market, thus necessitating a reliable system-level simulation tool for research and development. In this paper, we propose CXL-DMSim1, an open-source full-system simulator to simulate CXL disaggregated memory systems with high fidelity at a gem5-comparable simulation speed. CXL-DMSim incorporates a flexible CXL memory expander model along with its associated device driver, and CXL protocol support with CXL.io and CXL.mem. It can operate in both app-managed mode and kernel-managed mode, with the latter using a dedicated NUMA-compatible mechanism. The simulator has been rigorously verified against a real hardware testbed with both FPGA- and ASIC-based CXL memory devices, which demonstrates the qualification of CXL-DMSim in simulating the characteristics of various CXL memory devices at an average simulation error of 3.4%. The experimental results using LMbench and STREAM benchmarks suggest that the CXL-FPGA memory exhibits a ~2.88× higher latency than local DDR while the CXL-ASIC latency is ~2.18×; CXL-FPGA achieves 45-69% of local DDR memory bandwidth, whereas the number for CXL-ASIC is 82-83%. The study also reveals that CXL memory can significantly enhance the performance of memory-intensive applications, improved by 23× at most with limited local memory for Viper key–value database and approximately 60% in memory-bandwidth-sensitive scenarios such as MERCI. Moreover, the simulator’s observability and expandability are showcased with detailed case-studies, highlighting its great potential for research on future CXL-interconnected hybrid memory pool.
Yanjing Wang 0007, Lizhou Wu, Wentao Hong, Zicong Wang, Sunfeng Gao, Jie Zhang 0048, Sheng Ma, Dezun Dong, Xingyun Qi, Nong Xiao 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.10
2025 FPAMM: Fine-Grained Pipeline Architecture Accelerator for the Novel Transformer Architecture - Monarch Mixer
Hanyuan Li, Xingyun Qi, Puguang Liu, Zhenqi Li
ICA3PP (2)3
2025 Sumeru: An Efficient Hybrid-Granularity Cache Management Scheme for CXL-SSDs
Xuchao Xie, Qiulin Wu, Xingyun Qi, Zhenlong Song
ICA3PP (1)5
2025 PSCA: A FPGA-based Protein Structure Comparison Accelerator with Symmetric Simplified Matrix
Hui Su, Xingyun Qi, Qiang Wang 0006, Puguang Liu, Haoyu Liao
ICA3PP (6)3
2025 Enhancing Transformer Inference Efficiency on FPGA Through Fully Fusion and Integer-Only Quantization Techniques
abstract
The Transformer architecture has revolutionized the field of natural language processing (NLP) through its selfattention mechanism. However, its high computational complexity and memory requirement present significant deployment challenges on resource-constrained edge devices. While existing research predominantly focuses on accelerating linear operations via model compression and approximation techniques, the inefficiencies and high deployment costs of nonlinear operations (e.g., Softmax and LayerNorm) remain critically understudied. Although some studies have attempted to mitigate these challenges through techniques such as kernel fusion and integer-only quantization, these approaches still suffer from partial fusion and inefficient quantization with retained division operations, leaving significant efficiency gains unexploited. To bridge these gaps, we propose a fully fused Transformer accelerator that co-optimizes both linear and nonlinear operations while minimizing memory bottlenecks. For linear computations, our design incorporates a deeply optimized compute engine featuring double buffering, an output-stationary tiling strategy, and DSP-packing technology to maximize throughput. For nonlinear operations, we introduce a delayed computation strategy for vector-wise operators, effectively reducing memory bandwidth pressure and dependency stalls. Furthermore, we propose a hardware-efficient, divisionfree integer-only quantization scheme, leveraging$\log 2$quantization for Softmax and a polynomial-enhanced approximation for LayerNorm to eliminate costly floating-point units, thereby significantly reducing latency and resource overhead. Through systematic design space exploration, our solution, deployed on the Zynq Z-7100 platform, achieves 1.376 TOPS for BERT inference, demonstrating a$\mathbf{1. 6 5 - 2. 5 1} \boldsymbol{\times}$higher computational efficiency compared to prior works.
Zhenqi Li, Puguang Liu, Qiang Wang 0006, Yankang Zhao, Hanyuan Li, Xingyun Qi
ICCD8
2025 A Novel High-Speed Adaptive Duobinary Digital Detector Based on the Feed-Forward Equalizer and the Maximum Likelihood Sequence Detector for Wireline Transceivers
abstract
To solve the high bit error rate (BER) problem of conventional 56-Gb/s nonreturn-to-zero (NRZ) transceivers under high-insertion loss (IL) channels, this study proposes a high-speed adaptive duobinary (DB) digital detector based on the feed-forward equalizer (FFE) and the maximum likelihood sequence detector (MLSD). In this detector, adaptive FFE is combined with channel characteristics to generate DB signals and complete equalization, thus extending the transmission bandwidth and eye height and allowing a larger sampling phase offset. The parallel MLSD is used to complete the detection and decoding of DB signals to reduce the BER. An adaptive algorithm is proposed to avoid the long convergence time of the conventional zero-forcing (ZF) algorithm applied to the DB detector, so that it can be applied to various bit rates and IL channels. In this study, the verification of this DB detector is accomplished at 56 Gb/s. The platform based on a 56-Gb/s analog front-end chip (AFEC) and field-programmable gate array (FPGA) proves that the detector can work well in 12–56 Gb/s and multiple IL channels. The BER was less than 2e-8 at 56 Gb/s on −42-dB channel loss at 28 GHz. The structure can be well used for higher rate transceivers, such as 112 Gb/s.
Chaolong Xu, Fangxu Lv, Xingyun Qi, Qiang Wang 0006, Zhang Luo, Shijie Li 0002, Geng Zhang 0001
IEEE Trans. Very Large Scale Integr. Syst.4
2024 Optimization of TDM Using Single-ended Transmission for Multi-FPGA Platforms
abstract
In large-scale designs, the multi-FPGA system is a popular approach for hardware acceleration and pre-silicon verification due to its scalable capabilities. Researchers aim to enhance communication bandwidth and decrease latency between FPGAs, all while working within the constraints of limited physical pins available. To address this issue, we propose an optimization for time-division multiplexing(TDM). This optimization combines the advantages of traditional logic multiplexer circuits and high-speed serial circuits by utilizing the serial-to-parallel converter (ISERDES) and parallel-to-serial converter (OSERDES). The I/OSERDES approach typically employs the low-voltage differential signaling standard (LVDS) for high-speed data transmission, necessitating two physical pins. To save one physical pin, we advocate for a single-ended transmission solution without a forward clock. Moreover, we propose a new algorithm at the receiver to align the phase of the receiver’s clock. In comparison to the LVDS-based solution, our proposed interface achieves double the communication bandwidth of inter-FPGA chips without introducing additional system latency under the same TDM rate.
Haoyu Liao, Puguang Liu, Xingyun Qi
ISCAS6
2024 A High-performance Hardware Accelerator for Genome Alignment
abstract
Genome alignment is a vital process in genome sequencing and bio-informatics research. It entails aligning short DNA sequence fragments (usually tens to hundreds of base pairs) with a reference genome sequence, uncovering crucial biological information and variations. However, due to the high computational complexity of current sequence matching algorithms and the rapid growth of genetic data, there exists a computational bottleneck in sequence alignment workflows. Therefore, scholars have turned to using hardware to accelerate this computation, with a focus on high-performance computing. While existing acceleration solutions have mostly concentrated on the classical algorithm, showing significant improvements, there has been limited work on accelerating alignment algorithms in popular bio-informatics software tools (enhanced version). In this paper, we proposed a hardware accelerator for the alignment tasks in the popular sequence alignment tool minimap2 using the KSW2 algorithm. Our approach utilizes an anti-diagonal processing element(PE) array for parallel computation, implements the BAND technology in hardware, and enables support for longer input sequences without sacrificing accuracy. Additionally, we integrated the traceback stage of the KSW2 algorithm in hardware, reducing communication data overhead. We attained a 15.32x acceleration compared to software optimized with Streaming SIMD Extensions (SSE) instruction sets and hyper-threading technology. Additionally, our design shows a 3.42x speedup compared to other high-performance hardware.
Haoyu Liao, Hui Su, Xingyun Qi
ISPA7
2024 Automatic Implementation of Large-Scale CNNs on FPGA Cluster Based on HLS4ML
abstract
Convolutional Neural Networks (CNNs) have demonstrated remarkable performance across various computer vision tasks. Due to the computational and data-intensive nature of CNNs, Field-Programmable Gate Arrays (FPGAs) are exceptionally well-suited for accelerating the CNN computation process. However, large-scale CNNs such as ResNet-84 contain an enormous number of parameters that exceed the capacity of a single FPGA, rendering the deployment on a single device impractical. In this paper, we develop an automated end-to-end design flow for mapping large-scale CNNs across multiple FPGAs, on the basis of the HLS4ML dataflow architecture. We propose a graph optimization method to streamline the CNN structure and reduce resource consumption. We also summarize a resource allocation algorithm that automatically determines the specific hardware resources necessary for each CNN layer. Furthermore, we introduce a partitioning methodology capable of effectively segmenting CNNs into multiple FPGAs, and providing each subgraph with specific interfaces to support communication between different FPGAs. To validate the methodology, we construct a multi-FPGA platform interconnected via LVDS. We select two typical networks, ResNet-8 and ResNet-84, as the benchmarks for evaluation. The experimental results demonstrate that our approach significantly outperforms existing solutions. It attains an 18.6-fold increase in speed over Vitis AI, a 2.2-fold improvement over FINN, and a 3.4-fold enhancement over the original single-FPGA HLS4ML implementation for ResNet-8. For ResNet-84, our method achieves a remarkable 33.6-fold speedup over Vitis AI. Additionally, when compared to other non-automated multi-FPGA solutions, our methodology still exhibits significant performance improvements.
Xingyun Qi, Yankang Zhao, Zhenqi Li, Hanyuan Li, Qiang Wang 0006
ISPA2
2021 PFT: A Congestion Avoidance Method based on Proactive Flow Throttling at Endpoints
Xingyun Qi, Dezun Dong, Junsheng Chang, Jijun Cao
IM1
2021 MPICC: Multi-Path INT-Based Congestion Control in Datacenter Networks
Guoyuan Yuan, Dezun Dong, Xingyun Qi, Baokang Zhao
NPC3
2017 A Scalable and Resilient Microarchitecture Based on Multiport Binding for High-Radix Router Design
abstract
High-radix routers with low latency and high bandwidth play an increasingly important role in the design of large-scale interconnection networks such as those used in super-computers and datacenters. The tile-based crossbar approach partitions a single large crossbar into many small tiles and can considerably reduce the complexity of arbitration while providing throughput higher than the conventional switch implementation. However, it is not scalable due to power consumption, placement, and routing problems. In this paper, we propose a truly scalable router microarchitecture called Multiport Binding Tile-based Router (MBTR). By aggregating multiple physical ports into a single tile a high-radix router can be flexibly organized into a different array of tiles, thus the number of tiles and hardware overhead can be considerably reduced. Compared with a hierarchical crossbar, MBTR achieves up to 50%~75% reduction in memory consumption as well as wire area. Simulation results demonstrate MBTR is indistinguishable from the YARC router in terms of throughput and delay, and can even outperform it by reducing potential contention for output ports. We have fabricated an ASIC MBTR chip with 28nm technology. Internally, it runs at 700MHz and 30ns latency without any speedup. We also discuss how the microarchitecture parameters of MBTR can be adjusted based on the power, area, and design complexity constraints of the arbitration logic.
Kefei Wang, Gang Qu 0001, Liquan Xiao, Dezun Dong, Xingyun Qi
IPDPS6
2009 BOIN: A novel Bufferless Optical Interconnection Network for high performance computer
abstract
Most of the present optical interconnect networks within the high performance computers require either the buffering and opto-electonic conversion of the data packets or the pre-assigning of the optical path from the source to destination, which to a certain extent influence the performance metrics such as latency and throughput. Aiming at the limitation mentioned above, a hybrid optical-electrical interconnection for high performance computer system named bufferless optical interconnection network (BOIN) is brought forward together with the link control protocol and deadlock/livelock free routing algorithms. The upper bound of the data transmission latency within BOIN is also given. Experimental simulation is based on the comparison of BOIN and the other two similar networks. The results show that BOIN has advantages over the other two that it can deliver high throughput at low latency, which can well satisfy the need for high performance computing systems.
Xingyun Qi, Wei Yang 0037, Yongran Chen, Qiang Dou, Quanyou Feng, Wenhua Dou
AICCSA1