Xingqi Zou

dblp:241/8483 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
9since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 6 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 RTLMarker: Protecting LLM-Generated RTL Copyright via a Hardware Watermarking Framework
abstract
Recent advances of large language models in the field of Verilog generation have raised several ethical and security concerns, such as code copyright protection and dissemination of malicious code. Researchers have employed watermarking techniques to identify codes generated by large language models. However, the existing watermarking works fail to protect RTL code copyright due to the significant syntactic and semantic differences between RTL code and software code in languages such as Python. This paper proposes a hardware watermarking framework RTLMarker that embeds watermarks into RTL code and deeper into the synthesized netlist. We propose a set of rule-based Verilog code transformations, ensuring the watermarked RTL code's syntactic and semantic correctness. In addition, we consider an inherent tradeoff between watermark transparency and watermark effectiveness and jointly optimize them. The results demonstrate RTLMarker's superiority over the baseline in RTL code watermarking.
Kun Wang 0055, Mengdi Wang 0004, Xingqi Zou, Yinhe Han 0001, Ying Wang 0001
ASP-DAC4
2025 PC4: Precision Collective Communication Congestion Control for AI Cluster
Taoran Qi, Xingqi Zou, Liangce Deng, Guodong Wei
INFOCOM3
2025 ProMiner: Enhancing Locality, Parallelism, and Offloading for Graph Mining on Processing-in-Memory Systems
abstract
Graph mining, critical for discovering specific patterns within complex structures, is becoming increasingly important in our data-driven world. Due to their memory-bound nature, graph mining applications encounter significant limitations with conventional processor-centric systems, like central processing units (CPUs) and graphics processing units (GPUs), stemming from the costly data movement between memory and processing units. Memory-centric computing systems, such as processing-in-memory (PIM) where computation occurs directly within or near memory modules, have the potential to accelerate graph mining. However, accelerating graph mining applications with PIM presents three primary challenges: (1) the difficulty in utilizing locality, (2) the challenge of exploring parallelism, and (3) the complexity of workload offloading between PIM and CPU. Addressing these intricate challenges, we introduce ProMiner, a novel framework that integrates three key techniques through cohesive software and hardware co-design. First, we propose a partitioning method tailored for graph mining to enhance data locality. Second, we design a coarse-fine parallelism optimization scheme to explore parallelism across different levels of memory. Third, we introduce a concurrency-aware mechanism for performance estimation, aimed at identifying the optimal computing engine for workload offloading to maximize performance. Our experimental results demonstrate that ProMiner significantly advances the state-of-the-art in graph mining, achieving 48.8% and 29.9% execution time reduction over NDMiner and DIM- Mining, respectively.
Xiaoyang Lu, Xiaoming Chen 0003, Xingqi Zou, Yinhe Han 0001, Xian-He Sun
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2024 AceMiner: Accelerating Graph Pattern Matching using PIM with Optimized Cache System
abstract
Graph pattern matching (GPM), a critical algorithm for discovering specific patterns within complex structures, is becoming increasingly important in the data-driven world. GPM applications are memory-bound and can be accelerated by memory-centric computing systems, such as processing-in-memory (PIM). However, there are three primary challenges when it comes to accelerating GPM applications with PIM: (1) difficulty in utilizing locality, (2) heavy data movement, and (3) heavy comparison overhead due to pruning. To address these challenges, we propose AceMiner, a framework to accelerate GPM applications with a software and hardware co-design per-spective using PIM. In AceMiner, we embed hybridCache, a novel in-DRAM cache system with lower access latency and optimized replacement policy, to leverage the potential locality and reduce data movement in PIM. Additionally, we introduce a comparison unit to address the huge pruning overhead. Experimental results show that AceMiner outperforms the state-of-the-art, achieving speedups of 40.2% and 13.3% over NDMiner and DIMMining respectively, with less energy consumption and design overhead.
Xiaoyang Lu, Xiaoming Chen 0003, Xingqi Zou, Yinhe Han 0001, Xian-He Sun
ICCD5
2023 DrPIM: An Adaptive and Less-blocking Data Replication Framework for Processing-in-Memory Architecture
abstract
Processing-in-Memory (PIM) architecture suffers from frequent data synchronization, which raises inevitable coherent stalls and consumes more energy. This paper presents DrPIM, an adaptive and less-blocking data replication framework for Processing-in-Memory architecture to address such issues. The key insights are 1) finer-grained data management for the conflict data region and 2) the automatic generation and invalidations of data replications. The above two schemes significantly decrease the critical path of data synchronization and the access latency for conflict data, thus providing a less-blocking execution. Evaluations show that DrPIM achieves a speedup of 1.5x over the state-of-the-art and reduces the data-moving energy by 49%.
Hongyu Xue, Le Luo 0002, Xingqi Zou
ACM Great Lakes Symposium on VLSI5
2021 Analysing the Risk Propagation in the Project Portfolio Network using the SIRF Model
Xingqi Zou, Qing Yang 0026, Qinru Wang
ICORES1
2021 CoPIM: A Concurrency-aware PIM Workload Offloading Architecture for Graph Applications
abstract
Processing-in-Memory (PIM) is considered a promising solution to improve the performance of graph-computing applications by minimizing the data movement between the host and memory. Which workload to offload and how to offload it to PIM logic determine whether the PIM architecture is well utilized. Offloading too much or too little workload from the host processor to the PIM side could hurt overall performance. On the other hand, the offloading granularity needs to be representative without losing generality. In this paper, we present CoPIM, a novel PIM workload offloading architecture that can dynamically determine which portion of the graph workload can benefit more from PIM-side computation. CoPIM focuses on the loop code blocks of graph applications and evaluates the necessity of offloading based on a concurrent memory access model. We also provide detailed architectural designs to support the offloading. In this way, CoPIM reduces the size of offloading instructions and also improves the overall performance with less energy consumption. The experimental results show that compared with other state-of-the-art PIM workload offloading frameworks, CoPIM achieves a speedup by the geometric mean of 19.5% and 11.4% than PEI and GraphPIM, respectively. On the other hand, CoPIM also reduces the un-core energy consumption by 6.8% and 6.5% on average over PEI and GraphPIM, respectively.
Mingzhe Zhang 0005, Rujia Wang, Xiaoming Chen 0003, Xingqi Zou, Xiaoyang Lu, Yinhe Han 0001, Xian-He Sun
ISLPED5
2021 Breaking the von Neumann bottleneck: architecture-level processing-in-memory technology
Xingqi Zou, Xiaoming Chen 0003, Yinhe Han 0001
Sci. China Inf. Sci.1
2021 Dadu-Eye: A 5.3 TOPS/W, 30 fps/1080p High Accuracy Stereo Vision Accelerator
abstract
Stereo vision is widely deployed on robots and drones to enable depth estimation at a low cost. The combination of lightweight deep neural network (DNN) and cost volumes algorithm is proved to possess the advantages of both high depth estimation accuracy and speed. However, currently there is no accelerator architecture compatible with both efficient DNN inference and cost generation algorithms such as stereo matching. This work proposes a stereo vision accelerator called Dadu-eye, dedicated to real-time processing of high-resolution image streams. The proposed architecture adopts a pipelined hardware design with the techniques of operation approximation and scheduling-level optimization. First, a cost estimation block is designed to generate cost volumes from both luminance and color information. Second, a super pipelined multiplication and accumulation array with a row scan-based fused-layer convolution scheduling is proposed to perform the encoding and decoding neural network efficiently. Finally, an optical flow block is designed and cooperates with the array to approximately predict half of the frames’ depth to achieve real-time (30fps) processing on 1080p view. Based on the SMIC 40 nm CMOS process, this stereo vision accelerator achieves 5.3 TOPS/W power efficiency and significantly reduces 81% off-chip memory access.
Feng Min, Ying Wang 0001, Xingqi Zou, Yinhe Han 0001
IEEE Trans. Circuits Syst. I Regul. Pap.6
2019 Project Portfolio Risk Prediction and Analysis using the Random Walk Method
Xingqi Zou, Qing Yang 0026
ICORES1