EDBT 2026 Demo / reviewers in the wild / expert
Xuqi Zhu
dblp:84/8380
· DBLP profile ↗
10ranked-venue papers
4as first author
6since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Suika: Efficient and High-quality Re-scheduling of 3D-parallelized LLM Training Jobs in Shared ClustersabstractLarge Language Models (LLMs) are usually trained with 3D (data, tensor, and pipeline) parallelism—in shared GPU clusters where the available resources are highly dynamic. Rescheduling the idle resources to ongoing jobs can help improve cluster utilization, but doing so for 3D-parallelized training jobs suffers large overheads in performance modeling, decision making, and redeployment. We present Suika, a cluster training system that supports efficient and high-quality resource rescheduling for 3D-parallelized LLM training jobs. Suika holistically addresses the complexity challenges by exploiting the incremental nature of rescheduling. For performance modeling, it builds an accurate performance estimator with non-disruptive online profiling. For decision-making, it employs topology-aware sorting and an expand-and-balance algorithm to reduce the complexity of resource allocation and job parallelization, without compromising decision quality. Suika further integrates a device-to-device redeployment method to leverage the overlapping nature of incremental reconfiguration for overhead reduction. Experiments on 64-GPU physical cluster and 1024-GPU simulated cluster show that, Suika achieves 1.29 ~ 1.31× reduction in average JCT compared to state-of-the-art schedulers. Chen Chen 0067, Chunyu Xue, Qizhen Weng 0001, Zeren Li, Xuqi Zhu, Yongqiang Yang, Quan Chen 0002, Minyi Guo |
EuroSys | 8 |
| 2026 | Mitigating Scalability Challenges in LUT-Based Neural Networks via Pruning OptimisationsabstractModern deep neural networks heavily rely on a large number of multiply-accumulate operations, which constitute the predominant computational cost. To address this, Look-Up Table (LUT)-based matrix multiplications have emerged as a promising alternative for reducing the computational cost and time of the multiply-accumulate operations in a neural network. However, the LUT-based neural network still faces the scalability challenge due to the inherent limitations of LUT-based matrix multiplication. To mitigate these scalability limitations, this paper proposes a scalable and energy-efficient LUT-based approximate matrix multiplication unit (LUT-MU) constituting the basic component of the neural networks by integrating a pruning strategy on the MADDNESS algorithm, a LUT-based matrix multiplication methodology. With increasing problem size and precision demands in matrix multiplication, our proposed LUT-MU architecture effectively constrains resource expansion. The case study shows that deploying our LUT-MU in neural network architectures, including fully connected layers (MNIST) and ResNets (CIFAR-10, ImageNet)—on XCZU7EV and XCZU19EG FPGAs, produces up to 1.6× throughput improvement and 4.2× energy efficiency gains over mainstream CUDA-based network implementations, and 1.8× energy efficiency compared to leading quantised neural network implementations, with moderate impact on accuracy. Compared to original MADDNESS-based neural networks, our LUT-MU shows 1.3 to 2.6× resource savings based on various resolution configuration settings of MADDNESS. Xuqi Zhu, Huaizhi Zhang, Chandrajit Pal, Sangeet Saha, Klaus D. McDonald-Maier, Xiaojun Zhai |
IEEE Trans. Computers | 1 |
| 2025 | Late Breaking Results: Approximated LUT-Based Neural Networks for FPGA Accelerated InferenceabstractThis work presents LUT-MU, an approximated LUT-based Matrix Multiplication (MM) architecture designed for FPGA-based Neural Network (NN) inference across. The proposed architecture maximises the utilisation of on-chip memory bandwidth through dedicated memory distribution and pipeline design, addressing performance limitations inherent to LUT-based MM. Experimental evaluation demonstrates that LUT-MU achieves a four-fold improvement in NN inference throughput whilst reducing hardware resource consumption by 80% with only a 5% decline in accuracy. These results validate that our optimisation approach successfully resolves the performance constraints caused by the limited arithmetic intensity and memory bandwidth, enabling the LUT-MU to serve as a foundation for efficient NN acceleration systems. Xuqi Zhu, Huaizhi Zhang, Tamim M. Al-Hasan, Klaus D. McDonald-Maier, Xiaojun Zhai |
DATE | 1 |
| 2025 | Computationally Efficient FPGA-based Large Language Model Inference for Real-Time Decision-Making in Robotic SystemsabstractIntegrating Large Language Models (LLMs) into modern robotic systems presents significant computational and energy constraint challenges, particularly for human-centered robotic applications. This paper presents a novel hardware optimization technique for deploying LLMs on resource-constrained embedded devices, achieving an up to 77% reduction in computational latency through an FPGA implementation in comparison to other popular embedded computing devices (e.g., CPU and GPUs). Additionally, we demonstrate our methodology by deploying a LLaMA 2-7B model on a Unitree Go2 robotic dog integrated with the proposed FPGA platform. The proposed optimization framework preserves real-time interaction capabilities while significantly reducing computational and energy overhead, facilitating efficient natural language processing for human-robot interaction in safety-critical and dynamic environments. Experimental results demonstrate that the FPGA-based LLaMA 2-7B implementation achieves up to 6.06-fold and 1.95-fold higher throughput compared to baseline CPU and GPU implementations while maintaining comparable inference accuracy. Furthermore, the proposed FPGA design surpasses existing state-of-the-art FPGA implementations, delivering a 30% improvement in computational efficiency. Huaizhi Zhang, Tamim M. Al-Hasan, Xuqi Zhu, Weiyong Si, Klaus D. McDonald-Maier, Xiaojun Zhai |
IROS | 3 |
| 2025 | A Privacy-aware Quantilisation Approach for Efficient Edge Deep Learning AcceleratorabstractData privacy is one of the key concerns in machine learning model applications at the edge, especially in sensitive domains such as the healthcare sector. Here adversaries may exploit Membership Inference Attacks (MIAs) to determine if particular data points were used as part of datasets from the model’s training set, potentially leading to further data leakage issues. Although privacy preservation techniques like differential privacy (DP) can mitigate such risks during the training phase, this often results in degradation of model accuracy, making them less suitable for cloud training and edge deployment paradigms. For AI edge applications, existing research works for designing neural network accelerators primarily prioritize computational performance and power efficiency as their design target. In this paper, we introduce a novel privacy-aware quantilisation approach for deep learning accelerators and analyse the tradeoff between computational efficiency and privacy protection. The proposed system allows to adjust privacy constraints through tunable parameters, enabling flexible deployment on edge devices while meeting privacy and performance constraints. We have evaluated the proposed design on an AMD VCK190 board using a range of hypothetical MIA benchmarks. The results demonstrate that the proposed approach can effectively reduce the success rate of MIA attacks across multiple performance metrics. Huaizhi Zhang, Xuqi Zhu, Klaus D. McDonald-Maier, Xiaojun Zhai |
ISCAS | 2 |
| 2023 | Bayesian Optimization for Efficient Heterogeneous MPSoC Based DNN Accelerator Runtime TuningabstractWith the explosive growth of Internet of Things (IoT) devices and applications, deploying Deep Neural Networks (DNNs) on resource-constrained embedded edge devices has become a popular research trend. Because such systems have limited resources, they need to rely on optimising resource utilisation to meet performance requirements. However, for scenarios where the DNN application and workloads are dynamically changing, the offline system optimisation technique cannot achieve optimal runtime performance in practical environments. Hence, in this PhD project, we propose a Bayesian Optimisation (BO)-based runtime tuning scheme for improving energy efficiency of heterogeneous MPSoC-based DNN accelerator in the context of DNN applications. By seeking suitable hardware configurations of the accelerator for dynamic DNN inference workloads ranging from 200 M to 600 M FLOPs (floating-point operations) at runtime, the recommended configuration can averagely save up to 15.33% energy consumption from a random configuration setting. Xuqi Zhu, Sangeet Saha, Xiaojun Zhai, Klaus D. McDonald-Maier |
FPL | 1 |
| 2013 | Scalable Distributed Video Coding Using Compressed Sensing in Wavelet DomainabstractScalable video coding technologies provide adaptive video applications in heterogeneous and polytropical conditions. However, the highly hierarchical feature makes the loss or unsuccessful recovery of the base-layer to be catastrophic. In this paper, a distributed scalable video coding scheme using the new advances of compressed sensing is proposed to solve this problem at the energy-constrained encoder. The wavelet coefficients of the video frames have inherent fidelity scalability which is utilized in our scheme. Furthermore, the democracy of measurements reduces the risk of base-layer loss over the packet loss channel. Experimental results show that our scheme outperforms the existing scalable compressed sensing scheme by about 2dB. Nianfei Fan, Xuqi Zhu, Yu Liu 0001, Lin Zhang 0013 |
VTC Fall | 2 |
| 2011 | Distributed compressive video sensing based on smoothed ℓ0 norm with partially known supportabstractDistributed compressive video sensing (DCVS), aiming at capturing and compressing video data simultaneously, is an emerging field which exploits both intra- and inter-frame correlation. In this paper, we present a new algorithm based on smoothed ℓ0norm (SL0) which tries to directly minimize the ℓ0norm to decode a Wyner-Ziv frame when parts of its correlated key frame's support is known as side information (SI) in a typical DCVS scenario. With the assistance of the modified initialization, our proposed algorithm can reconstruct the Wyner-Ziv frame of the same accuracy with much lower measurement rate compared to the case when the partially known support is not used as SI. It is experimentally shown that our proposed scheme outperforms GPSR at the expense of a tolerable decoding complexity. When compared with modified-cs, a large saving in decoding CPU time is achieved in sacrifice of some PSNR performance. Yu Liu 0001, Lin Zhang 0013, Xuqi Zhu |
ICME | 4 |
| 2011 | An unequally protected Distributed Compressed Video Sensing algorithmabstractDistributed Compressed Video Sensing (DCVS) has developed as one of the efficient solutions that guarantee low complexity video compression. In this paper, a novel DCVS algorithm with unequal protection of the video signal's elements is proposed. The new algorithm utilizes not only the sparsity and probability distribution of the video signal but also its particular unequal significance feature. Based on this feature, we design the structured irregular low-density sensing matrix to sample the signal. From the analysis and simulation results, it is confirmed that our method has higher recovery quality than the conventional Bayesian Compressed Sensing (CS) using Belief Propagation (BP). Moreover, the excellent noise-resilience of BP is preserved in our algorithm comparing to the DCVS schemes using optimization recovery. Bin Li 0022, Xuqi Zhu, Yu Liu 0001, Lin Zhang 0013 |
VCIP | 2 |
| 2010 | Symmetric Distributed Joint Source-Channel Coding Using Raptor CodesabstractDistributed compression of correlated sources has been discussed much in wireless sensor networks, while the error-resilient implementation of this efficient coding strategy is one of the crucial issues for applications. In this paper, a symmetric Distributed Joint Source-Channel Coding (DJSCC) scheme is proposed by using Raptor codes for the independent channels case. The channel noise and the correlation of sources are considered simultaneously within one set of encoder and decoder. The symmetric structure of the proposed approach leads to more flexible and balanced rate allocation, and the rateless property of Raptor codes also guarantees tractable code rates and error correction. At last, the simulation results demonstrate that our scheme outperforms the existing LDPC-based scheme at low SNR. Xuqi Zhu |
NAS | 1 |