Xiaolong Lin

dblp:229/3306 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Privacy-Preserving Sketches for Securely Estimating Intersection Cardinality Over Distributed Data Sets
abstract
Computing the number of distinct elements (i.e., cardinality) in the intersection of two sets is a fundamental task in various distributed systems, including measuring origin-destination flows in wide-area networks and data synchronization in distributed databases. Due to the enormous data scale, lightweight probabilistic methods, such as FM sketch and HyperLogLog sketch, are extensively used in these systems to estimate the set intersection cardinality, with memory efficiency, high accuracy, and low communication costs. However, if a set's sketch and the hash functions used to construct the sketch are disclosed to an untrusted third party, the privacy of the set's sensitive elements may be compromised. Applying the differential privacy mechanism directly to safeguard the sketch's privacy may incur significant estimation errors. To address this challenge, we propose a novel private sketch,SetXor, for securely estimating intersection cardinality for static sets. Specifically, we incorporate the randomized response noise into the constructed sketch to achieve local differential privacy while ensuring ourSetXorsketch is mergeable. We establish a concrete probabilistic model to mitigate the estimation error caused by the noise and theoretically analyze the variance. We further propose a novel sketchSetXorDynenabling intersection cardinality estimation for streaming sets where elements appear sequentially and contain duplicates. We employ a sampling-like method to eliminate the impact of different parities of element occurrences, allowing us to handle all elements without bias. We conduct extensive experiments on synthetic and four real-world datasets. The results demonstrate that our methods reduce the Average Absolute Relative Error (AARE) of state-of-the-art baselines by up to$110\times$on synthetic datasets and$80\times$on real-world datasets, while achieving up to$18.7\times$speedup under the same settings.
Pinghui Wang, Zhicheng Li 0007, Xiaolong Lin, Rundong Li 0002
IEEE Trans. Dependable Secur. Comput.4
2025 BMP-SD: Marrying Binary and Mixed-Precision Quantization for Efficient Stable Diffusion Inference
abstract
Stable Diffusion (SD) is an emerging deep neural network (DNN) model that has demonstrated impressive capabilities in generative tasks such as text-to-image generation. However, the iterative denoising stage of the SD model is extremely expensive in both computations and memory accesses, making it challenging for fast and energy-efficient edge deployment. To alleviate the overhead of denoising, in this paper we propose BMP-SD, a post-training quantization framework for hardware-efficient SD inference. BMP-SD employs binary weight quantization to significantly reduce the computational complexity and memory footprint of iterative denoising, along with dynamic step-aware mixed-precision activation quantization, based on the observation that not all denoising steps are equally important for a specific input prompt. Experiments on the text-to-image generation task show that BMP-SD achieves mixed-precision (W1.73A4.87) with minimal accuracy loss on MS-COCO 2014 dataset. We also evaluate the BMP-SD quantized model on three state-of-the-art bit-flexible DNN accelerators, results reveal that our method can deliver up to 5.14× performance and 3.85×energy efficiency improvements compared to W8A8 quantization.
Xiaolong Lin, Jiayao Ling, Xiaoyao Liang
DATE3
2025 SBQ: Exploiting Significant Bits for Efficient and Accurate Post-Training DNN Quantization
abstract
Post-Training Quantization is an effective technique for deep neural network acceleration. However, as the bit-width decreases to 4 bits and below, PTQ faces significant challenges in preserving accuracy, especially for attention-based models like LLMs. The main issue lies in considerable clipping and rounding errors induced by the limited number of quantization levels and narrow data range in conventional low-precision quantization. In this paper, we present an efficient and accurate PTQ method that targets 4 bits and below through algorithm and architecture co-design. Our key idea is to dynamically extract a small portion of significant bit terms from high-precision operands to perform low-precision multiplications under the given computational budget. Specifically, we propose Significant-Bit Quantization (SBQ). It exploits a product-aware method to dynamically identify significant terms and an error-compensated computation scheme to minimize compute errors. We present a dedicated inference engine to unleash the power of SBQ. Experiments on CNNs, ViTs, and LLMs reveal that SBQ consistently outperforms prior PTQ methods under 2~4-bit quantization. We also compare the proposed inference engine with state-of-the-art bit-operation-based quantization architectures TQ and Sibia. Results show that SBQ can achieve the highest area and energy efficiency.
Jiayao Ling, Gang Li 0015, Qinghao Hu 0001, Xiaolong Lin, Jian Cheng 0001, Xiaoyao Liang
DATE4
2025 Light-DiT: An Importance-Aware Dynamic Compression Framework for Diffusion Transformers
Gang Li 0015, Xuan Zhang 0001, Jiayao Ling, Xiaolong Lin, Zhuoran Song, Jian Cheng 0001, Xiaoyao Liang
Euro-Par (2)5
2025 Transcript-Prompted Whisper with Dictionary-Enhanced Decoding for Japanese Speech Annotation
Xiaolong Lin, Jiawang Liu, Shixi Huang, Zhenpeng Zhan
INTERSPEECH2
2025 DualCodec: A Low-Frame-Rate, Semantically-Enhanced Neural Audio Codec for Speech Generation
Jiaqi Li 0030, Xiaolong Lin, Zhekai Li, Shixi Huang, Yuancheng Wang, Chaoren Wang, Zhenpeng Zhan, Zhizheng Wu 0001
INTERSPEECH2
2025 An Efficient Bit-Sparse DNN Accelerator Exploiting Adaptive Bit-Serial Computations
abstract
Bit sparsity, an intrinsic attribute of binary representation, has been widely utilized in DNN inference acceleration. Despite the advantages in performance and energy efficiency demonstrated by existing bit-serial-based bit-sparse accelerators, they still face two notable limitations: 1) At the low-level bit-serial multiplier level, existing methods either statically select weight or activation as the serialized object during the design phase, or simply serialize both without considering the distribution of non-zero bits in different operands, thereby failing to achieve optimal performance; 2) At the high-level dataflow level, existing approaches do not eliminate zero values in data movement and computation, leading to considerable energy and latency overhead, as well as suboptimal PE utilization. In this work, we propose AdaS-Pro accelerator for fast and energy-efficient DNN inference. At the multiplier level, AdaSPro employs an adaptive bit-serial computation scheme, which dynamically serializes the input operand with fewer non-zero bits at runtime, thereby minimizing compute cycles. To further enhance performance, AdaS-Pro introduces an improved Booth encoding method to reduce the number of non-zero bits in each operand. At the dataflow level, AdaS-Pro employs a compressed format to eliminate zero values and proposes a bi-directional inner-join unit coupled with a ring-shaped scheduler to achieve efficient non-zero workload extraction and balancing. Experimental results show that AdaS-Pro outperforms existing state-of-theart bit-sparse accelerators, such as BitLet, BitX, and Laconic, with performance improvements of 4.03×, 6.78×, and 1.43×, respectively.
Jiayao Ling, Gang Li 0015, Xiaolong Lin, Xing Li 0031, Jian Cheng 0001, Xiaoyao Liang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2024 Early: An Importance-Aware Early Firing and Exit for SNN Acceleration
abstract
Spiking neural networks (SNNs) have been promising applications in the image recognition domain, and their key component is the spiking neuron. SNN s mainly contain integration and firing processes, which are essentially weight accumulation and threshold comparison, respectively. However, spike trains of the neurons exhibit high sparsity and irregularity in both temporal and spatial domains, leading to inefficient memory access and computation. Therefore, designing an efficient accelerator for SNNs is urgent. This paper presents an elaborate accelerator Early in a software-hardware co-design way. At the software level: (i) Noticing the importance of weights, where larger weights disproportionately affect the membrane potential, we devise a weight importance-aware early firing solution for the firing neurons. It prioritizes the accumulation of these large weights, thereby accelerating the membrane potential's rise to surpass the threshold sooner. (ii) Meanwhile, given the observation that a large proportion of neurons do not eventually be fired even after experiencing a long delay of weight accumulation, we propose a weight importance-aware early exit mechanism. It preferentially accumulates large weights and compares the membrane potential with the predetermined threshold, which early halts the accumulation of neurons that are unlikely to be fired, enhancing efficiency. At the hardware level, we design a specialized processing element (PE) featuring the reorder engine for spikes and weights, tailored to realize the aforementioned strategies. Experimental results show that Early averagely achieves 20.3 x, 6.5 x, and 2.4 x speedup compared to the state-of-the-art accelerators Spinalflow, PTB, and SATO. Meanwhile, it averagely achieves 25.2x, 7.4x, and 3.2x energy savings with respect to the three accelerators.
Xuan Zhang 0001, Zhuoran Song, Peng Zhou 0030, Xing Li 0031, Xueyuan Liu 0001, Xiaolong Lin, Zhezhi He, Li Jiang 0002, Naifeng Jing, Xiaoyao Liang
ICCD6
2024 GNeRF: Accelerating Neural Radiance Fields Inference via Adaptive Sample Gating
abstract
NeRF is an emerging algorithm in computer graphics that has achieved state-of-the-art results in areas such as image rendering and 3D reconstruction. However, to compute the RGB of pixels in a view, NeRF executes MLP calculations on a huge number of sample points, resulting in significant computational complexity. To address this issue, we propose a simple and hardware-friendly NeRF algorithm (dubbed GNeRF) in this paper. GNeRF is designed based on the concept of "gating-by-decomposing". Specifically, It decomposes the original large MLP into two smaller branches. For each ray, GNeRF utilizes one branch to predict the important samples based on the direction information adaptively. The RGB calculations are then solely performed on these important samples using the other branch. Experimental results show that GNeRF can achieve comparable PSNR with only 3% FLOPS of the original NeRF. To showcase the hardware efficiency of GNeRF, we also design an FPGA-based NeRF accelerator on Xilinx ZCU102 MPSoC. Evaluation reveals that GNeRF can significantly enhance inference performance with minimal modifications to the existing MLP engine.
Gang Li 0015, Xiaolong Lin, Jiayao Ling, Xiaoyao Liang
ISCAS3
2023 AdaS: A Fast and Energy-Efficient CNN Accelerator Exploiting Bit-Sparsity
abstract
Bit-sparsity has shown its promise in CNN acceleration. However, prior bit-sparse accelerators have two drawbacks: 1) a large number of zero values are involved in the computation and data movement; 2) the distribution of non-zero bits is not considered in PE design. To address these issues, we propose AdaS. At the multiplier level, we dynamically serialize the operands that have fewer non-zero bits. At the dataflow level, we propose a group-wise bi-directional inner-join for workload extraction and balancing. Results show that AdaS can achieve 3.28×, 2.05× speedup, and 1.99×, 1.80× energy efficiency over Bit-Pragmatic and Laconic, respectively.
Xiaolong Lin, Gang Li 0015, Zizhao Liu, Zhuoran Song, Naifeng Jing, Xiaoyao Liang
DAC1
2020 Deep Representation Learning for Location-Based Recommendation
abstract
Location-based recommendation has recently received a lot of attention in the communities of information service and mobile application. Its task is to provide personalized recommendations of points of interest (POIs) to users at a certain time and location. However, existing location-based recommendation models have at least two main drawbacks: first they cannot adequately capture semantic features of POIs and users, which may lead to unsatisfactory recommendations and second they cannot effectively address the cold-start problem. To address the above drawbacks, in this article, we first propose a novel deep representation learning-based model (DRLM) for improving the recommendation accuracy. In DRLM, we mainly focus on learning to accurately represent semantic features of POIs and users. Specifically, four co-occurrence matrices are constructed to produce four different original features for each POI, and a principal component analysis (PCA) algorithm is utilized to generate a semantic feature of each POI from its four original features. On the other hand, a three-modal simple recurrent unit (TMSRU) network is given to constructed semantic features of users using semantic features of POIs, times, and locations. We further propose minimum description length (MDL)-based and skyline-based strategies to address the cold-start issues for new users and new POIs, respectively. Through experiments on two real-world data sets, we show that compared with the state-of-the-art approaches, the proposed model DRLM can achieve the superior performance in terms of high recommendation accuracy and effectiveness in handling the cold-start problem.
Zhenhua Huang 0001, Xiaolong Lin, Hai Liu 0006, Bo Zhang 0004, Yunwen Chen, Yong Tang 0001
IEEE Trans. Comput. Soc. Syst.2