Xinrui Zhu

dblp:245/3426 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Electronic design automation · 50% Hardware accelerators and domain-specific architectures · 50%
Artificial intelligence
1 paper
Reinforcement learning · 44% Deep learning architectures and training · 28% Trustworthy machine learning · 28%
Theoretical computer science
1 paper
Algorithms and data structures · 100%

Topics — the 12 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures › machine learning accelerator
CNN accelerator
0.912025
DSPlacer: DSP Placement for FPGA-based CNN Accelerator · DAC 2025
Hardware accelerators and domain-specific architectures › machine learning accelerator › CNN accelerator
FPGA-based CNN accelerator
0.912025
DSPlacer: DSP Placement for FPGA-based CNN Accelerator · DAC 2025
Electronic design automation › physical design › placement › circuit placement
FPGA placement
0.912025
DSPlacer: DSP Placement for FPGA-based CNN Accelerator · DAC 2025
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.912025
DSPlacer: DSP Placement for FPGA-based CNN Accelerator · DAC 2025
Electronic design automation
physical design
0.912025
DSPlacer: DSP Placement for FPGA-based CNN Accelerator · DAC 2025
Electronic design automation › physical design
placement
0.912025
DSPlacer: DSP Placement for FPGA-based CNN Accelerator · DAC 2025
Algorithms and data structures › sequence algorithms
string algorithms
0.912025
Double-ended palindromic trees in linear time · Inf. Comput. 2025
Machine learning › Deep learning architectures and training
mixture of experts
0.412019
Self-Supervised Mixture-of-Experts by Uncertainty Estimation · AAAI 2019
Machine learning › Reinforcement learning
multi-task reinforcement learning
0.412019
Self-Supervised Mixture-of-Experts by Uncertainty Estimation · AAAI 2019
Machine learning › Trustworthy machine learning
uncertainty estimation
0.412019
Self-Supervised Mixture-of-Experts by Uncertainty Estimation · AAAI 2019
Machine learning › Reinforcement learning › off-policy reinforcement learning
experience replay
0.112019
Self-Supervised Mixture-of-Experts by Uncertainty Estimation · AAAI 2019
Machine learning › Reinforcement learning
sample efficiency
0.112019
Self-Supervised Mixture-of-Experts by Uncertainty Estimation · AAAI 2019

Methods — techniques the papers use, named apart from their topics

min-cost flow · 0.9integer linear programming · 0.9graph convolutional network · 0.9uncertainty estimation · 0.4mixture of experts · 0.4experience replay · 0.4deep deterministic policy gradient · 0.4
YearPublicationVenuePosition
2025 DSPlacer: DSP Placement for FPGA-based CNN Accelerator
abstract
Deploying convolutional neural networks (CNNs) on hardware platforms like Field Programmable Gate Arrays (FPGAs) has garnered significant attention due to their inherent flexibility and parallelism. Achieving optimal timing closure remains a critical challenge, as placement directly impacts clock frequency and throughput. Existing approaches often face scalability issues with large designs or fail to formalize placement rules into automated algorithms. In this paper, we propose DSPlacer, a novel DSP placement framework designed for diverse CNN accelerator architectures in the context of FPGA design. The proposed approach iteratively optimizes the placement of datapath DSPs to enhance timing performance. To achieve this, DSPlacer integrates several advanced techniques, including graph convolutional network-based datapath DSP identification, DSP graph construction, min-cost-flow DSP assignment, and integer linear programming (ILP)-based cascade constraint legalization. These techniques collectively address two key requirements for datapath DSP placement: (1) cascading datapath DSPs to achieve a compact layout, and (2) preserving direct datapath information between the processing system and programmable logic. The framework has been evaluated on multiple academic benchmarks and compared against AMD Xilinx Vivado 2020.2 and AMF-Placer 2.0. Experimental results demonstrate that DSPlacer improves Worst Negative Slack (WNS) by 32% and 65%, respectively, highlighting its efficacy and superiority.
Baohui Xie, Xinrui Zhu, Yuan Pu 0001, Tongkai Wu, Xiaofeng Zou, Bei Yu 0001, Tinghuan Chen
DAC2
2025 Natural Language to Overpass Query: A Multi-Step Approach Using Task Decomposition and Key-Value Correction
abstract
We investigate the challenge of generating OverpassQL from natural language in the Text-to-OverpassQL task and explore the data in the existing OverpassNL dataset. To address the structural mismatch between natural language and OverpassQL, we propose a task decomposition-based multi-step prompting approach that generates auxiliary information to help align natural language with OverpassQL structures, thereby enhancing model performance. Furthermore, we introduce a Key-Value Correction Module specifically targeting key-value pair matching difficulties in Text-to-OverpassQL tasks, designed to rectify potential syntactic errors and key-value mismatches in generated queries. Our experiments on GPT-3.5 Turbo and GPT-4 demonstrate absolute performance gains of$\mathbf{1. 4 \%}$and 0.6 % respectively. Under retrieval-augmented setting ablation, we achieve a more significant 3.5 % improvement with GPT-3.5 Turbo. Experimental results confirm that our method consistently improves performance across various models and configurations, particularly showing enhanced effectiveness in medium and small-scale models.
Xinrui Zhu, Xuan Wang 0002, Yuanfeng Song, Hanlin Gu
MDM2
2025 Double-ended palindromic trees in linear time
abstract
The palindromic tree (a.k.a. eertree) is a data structure that provides access to all palindromic substrings of a string. In this paper, we propose a dynamic version of eertree, called double-ended eertree, which supports online operations on the stored string, including double-ended queue operations, counting distinct palindromic substrings, and finding the longest palindromic prefix/suffix. At the heart of our construction, we identify a new class of substring occurrences, called surfaces, that are palindromic substring occurrences that are neither prefixes nor suffixes of any other palindromic substring occurrences, which is of independent interest. Surfaces characterize the link structure of all palindromic substrings in the eertree, thereby allowing a linear-time implementation of double-ended eertrees through a linear-time maintenance of surfaces.
Qisheng Wang, Ming Yang 0033, Xinrui Zhu
Inf. Comput.3
2025 Distributed MIMO Radar Network for IoT: High-Resolution 4-D Point Cloud Generation and Signal Processing for Smart Mobility
abstract
The rapid advancement of Internet of Things (IoT) technology, particularly in the realm of autonomous driving, has elevated the requirements for automotive radar systems to achieve precise environmental perception. This article delves into the distributed MIMO radar network model and the associated signal processing techniques that enable accurate measurement of range, velocity, and angular positions, culminating in the generation of high-resolution 4-D point clouds. These capabilities are pivotal for intelligent interactions between vehicles and their surroundings within the IoT ecosystem. We introduce a stepped-frequency frequency-modulated continuous wave (SF-FMCW) waveform that incrementally increases the starting frequency of each chirp, leading to a larger bandwidth and finer range resolution without altering the individual chirp bandwidth. Furthermore, we propose a hybrid FDM-DDM scheme to ensure orthogonality among MIMO waveforms. This scheme allows for decoding across various DDM modes through the application of overlapping binary masks, while maintaining unambiguous range-Doppler measurements, which is crucial for real-time data processing and decision-making within IoT. To enhance angular resolution, we optimize the array configuration and develop a low sidelobe direction of arrival (DOA) estimation method using phase coherence factor (PCF) techniques. Extensive simulations and experimental analyses demonstrate the superior performance of the proposed methods in resolving closely spaced targets and generating high-fidelity 4-D point clouds, even in challenging scenarios with limited angular separation. The development of these technologies is significant for intelligent perception and safe navigation of vehicles within the IoT, providing a technological foundation for seamless integration of vehicles with the IoT infrastructure.
Yi Li 0066, Weijie Xia, Lingzhi Zhu, Cao Qu, Xinrui Zhu, Jianjiang Zhou
IEEE Internet Things J.5
2024 RISCSparse: Point Cloud Inference Engine on RISC-V Processor
abstract
Machine learning on point clouds is increasingly accessible at the edge, notably in applications such as autonomous driving. However, the sparse and irregular nature of point clouds presents significant latency challenges on general-purpose hardware. RISC-V, with its evolving ecosystem, offers a promising platform for embedding intelligence at the edge due to its full-stack scalability. This paper focuses on the advanced point cloud operation known as submanifold convolution (SC), deploying submanifold sparse convolutional networks (SSCNs) on a RISC-V System-on-Chip (SoC) designed within the Chipyard framework. We address three critical bottlenecks of SSCNs- Rule Map Construction (Mapping), Gather-MatMul-Scatter (GMS), and uncombined operation - to meet the real-time inference requirement for the on-chip implementation. By leveraging the RISC-V Vector extension and Gemmini, an open-source full-stack DNN accelerator generator, we vectorize the Mapping process, offload GEMM-related operations to the Gemmini Systolic Array, and cooperatively use the Systolic Array and vector processing units to reduce the memory footprint. Our evaluations show that the RISC-V-based SSCNs implementation achieves an average of 11.73× and 13.1× overall speedups with a small workload compared to TorchSparse on Edge-CPU, a state-of-the-art point cloud inference engine, for 3D segmentation and detection tasks, respectively. When contrasting with TorchSparse on an Edge-GPU, our implementation still delivers a notable improvement, with average speedups of 1.63× for 3D segmentation and 1.07× for detection tasks.
Shangran Lin, Xinrui Zhu, Baohui Xie, Tinghuan Chen, Cheng Zhuo, Qi Sun 0002, Bei Yu 0001
ICCAD2
2024 Key Flow First Prioritized Flow Scheduling Strategy in Multi-Tenant Data Centers
abstract
The mixed flow in multi-tenant data centers presents a challenge for priority flow scheduling due to the coexistence of various requirements such as latency and throughput. To address this issue, we propose Key Flow First (KFF), a balanced scheduling algorithm suitable for mixed flows in multi-tenant data centers. Firstly, KFF categorizes flows into Latency-Sensitive Flows (LS Flow) and Throughput-Demanding Flows (TD Flow) based on the Quality of Service (QoS) of their application sources. Secondly, it further differentiates flows into Mice Flows and Elephants Flows based on the amount of already sent bytes. Thirdly, KFF employs the Multi-Level Feedback Queue (MLFQ) threshold update algorithm and a priority-based strict forwarding mechanism. By avoiding reliance on complex flow priors, KFF consistently maintains reasonable scheduling of mixed flows under different load scenarios. Experimental results demonstrate that KFF effectively reduces the real-time load on the network and achieves good performance in terms of MAX (Shortest Job First (SJF), Earliest Deadline First (EDF)) performance under diverse load conditions. Compared to PIAS, KFF reduces the FCT slow down of deadline flows by nearly 60% under high TD loads; compared to Karuma and Time Deadline Aware pFabric (TDA-pFabric), KFF reduces the flow completion time (FCT) slow down of non-deadline Mice flows by over 90% under high LS loads and meanwhile guaranteeing nearly 0 deadline miss rate.
Xudong Tao, Xiaoyan Qian 0002, Weibei Fan, Yuzhou Shi, Xinrui Zhu, Shuwen Wei
IEEE Trans. Netw. Serv. Manag.6
2019 Self-Supervised Mixture-of-Experts by Uncertainty Estimation
abstract
Learning related tasks in various domains and transferring exploited knowledge to new situations is a significant challenge in Reinforcement Learning (RL). However, most RL algorithms are data inefficient and fail to generalize in complex environments, limiting their adaptability and applicability in multi-task scenarios. In this paper, we propose SelfSupervised Mixture-of-Experts (SUM), an effective algorithm driven by predictive uncertainty estimation for multitask RL. SUM utilizes a multi-head agent with shared parameters as experts to learn a series of related tasks simultaneously by Deep Deterministic Policy Gradient (DDPG). Each expert is extended by predictive uncertainty estimation on known and unknown states to enhance the Q-value evaluation capacity against overfitting and the overall generalization ability. These enable the agent to capture and diffuse the common knowledge across different tasks improving sample efficiency in each task and the effectiveness of expert scheduling across multiple tasks. Instead of task-specific design as common MoEs, a self-supervised gating network is adopted to determine a potential expert to handle each interaction from unseen environments and calibrated completely by the uncertainty feedback from the experts without explicit supervision. To alleviate the imbalanced expert utilization as the crux of MoE, optimization is accomplished via decayedmasked experience replay, which encourages both diversification and specialization of experts during different periods. We demonstrate that our approach learns faster and achieves better performance by efficient transfer and robust generalization, outperforming several related methods on extended OpenAI Gym’s MuJoCo multi-task environments.
Zhuobin Zheng, Chun Yuan 0003, Xinrui Zhu, Zhihui Lin, Yangyang Cheng, Jiahui Ye
AAAI3
2019 Image-to-Tree: A Tree-Structured Decoder for Image Captioning
abstract
Automatically generating natural language descriptions of images is a fundamental problem in artificial intelligence that connects computer vision and natural language processing. In recent years tremendous success has been shown in image captioning under the encoder-decoder framework, in which decoders are often chain-structured with Recurrent Neural Networks(RNNs), treating sentences as sequences. However, natural sentences are not inherently linear structures, but hierarchical structures. In this paper, we for the first time proposed a model with tree-structured decoder for image captioning(Image-to-Tree), which does not directly generate sentences but instead explicitly generates their dependency trees in a top-down manner. Inspired by the success of attention mechanism in image captioning, we also proposed a corresponding attention-based model for Image-to-Tree. Experiments on MSCOCO dataset demonstrate that our model can achieve comparable results to chain-structured models of different language metrics.
Zhiming Ma, Chun Yuan 0003, Yangyang Cheng, Xinrui Zhu
ICME4