Kang You

dblp:304/3304 · DBLP profile ↗
← Back
16ranked-venue papers
8as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 6 first-author · 9 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 ELSA: An Elastic Snn Inference Architecture for Efficient Neuromorphic Computing
Kang You, Chen Nie, Lee Jun Yan, Ziling Wei, Yu Feng 0007, Honglan Jiang, Zhezhi He
ISCA1
2026 APU: Accelerate Point Cloud Neural Networks via Unified Processing-in-SRAM Architecture
abstract
Recent advances in deep learning have expanded point cloud applications by point-based neural networks (PNNs). However, the escalating complexity and computational demands of PNNs overwhelm conventional computers. Specialized PNN accelerators have emerged, significantly outperforming modern CPUs and GPUs. Nevertheless, existing designs remain inefficient when handling performance-critical mapping kernels of PNNs, involving diverse arithmetic functions (e.g., add, multiply, sort) across separate hardware modules. This fragmentation restricts hardware sharing and data locality, leading to area overhead, redundant data movements, and under-utilization. Therefore, a unified and efficient micro-architecture for mapping kernels is needed to enhance performance and reduce data transfers. This paper presents APU, an efficient processing-in-memory (PIM) architecture for PNN acceleration. We introduce the first unified SRAM-PIM micro-architecture that supports all mapping kernels in mainstream PNNs. Data movement is reduced through extensive on-chip memory and maximized data locality viain-situcomputing approach. At the algorithmic level, we introduce mask grouping and aggregation to eliminate costly sorting operations, enabled by hardware support for in-memory vector max-search. This refined strategy reduces computational overhead and data transfers while improving inference accuracy.We further enhance performance by exploiting parallelism across PNN operations and applying mixed-precision quantization. Evaluated on real-world PNN workloads, APU outperforms the state-of-the-art accelerator by 2.54× in speedup and 4.54× in energy saving.
Chen Nie, Kang You, Yu Feng 0007, Limin Xiao 0002, Weifeng Zhang 0003, Zhezhi He
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2026 EDRIC: Embracing 1D Autoencoder for Real-Time Lossy LiDAR Reflectance Compression
abstract
While recent advancements in LiDAR reflectance compression have improved rate-distortion performance, real-time processing remains an unresolved challenge. In this work, we introduce EDRIC, a highly effective neural compression framework that offers state-of-the-art compression efficiency while achieving real-time capability. To overcome the suboptimal downsampling scheme and inefficient feature extraction in conventional 3D frameworks, EDRIC serializes a 3D point cloud into a 1D sequence and introduces a lightweight 1D autoencoder to efficiently compress the serialized LiDAR reflectance signal. In addition, we explicitly incorporate geometric priors through a geometry-aware entropy model, effectively exploiting the interdependencies between reflectance attributes and underlying geometry. Extensive experiments on representative datasets (e.g., KITTI, Ford, nuScenes, and QNX) demonstrate that EDRIC achieves 8.38%~9.85% BD-BR reduction compared to the latest G-PCCv23 (RAHT) standard while operating $22\times $ faster (e.g., >50 frames per second on an RTX 4090 GPU). Furthermore, EDRIC comprises merely 2.3M parameters, making it a practical solution for real-world deployment.
Kang You, Kequan Mao, Dandan Ding, Zhan Ma 0001
IEEE Trans. Image Process.2
2025 RENO: Real-Time Neural Compression for 3D LiDAR Point Clouds
abstract
Despite the substantial advancements demonstrated by learning-based neural models in the LiDAR Point Cloud Compression (LPCC) task, realizing real-time compression—an indispensable criterion for numerous industrial applications—remains a formidable challenge. This paper proposes RENO, the first real-time neural codec for 3D LiDAR point clouds, achieving superior performance with a lightweight model. RENO skips the octree construction and directly builds upon the multiscale sparse tensor representation. Instead of the multi-stage inferring, RENO devises sparse occupancy codes, which exploit cross-scale correlation and derive voxels’ occupancy in a one-shot manner, greatly saving processing time. Experimental results demonstrate that the proposed RENO achieves real-time coding speed, 10 fps at 14-bit depth on a desktop platform (e.g., one RTX 3090 GPU) for both encoding and decoding processes, while providing 12.25% and 48.34% bit-rate savings compared to G-PCCv23 and Draco, respectively, at a similar quality. RENO model size is merely 1MB, making it attractive for practical applications. The source code is available at https://github.com/NJUVISION/RENO.
Kang You, Tong Chen 0004, Dandan Ding, Muhammad Salman Asif, Zhan Ma 0001
CVPR1
2025 VISTREAM: Improving Computation Efficiency of Visual Streaming Perception via Law-of-Charge-Conservation Inspired Spiking Neural Network
abstract
Visual streaming perception (VSP) involves online intelligent processing of sequential frames captured by vision sensors, enabling real-time decision-making in applications such as autonomous driving, UAVs, and AR/VR. However, the computational efficiency of VSP on edge devices remains a challenge due to power constraints and the under-utilization of temporal dependencies between frames. While spiking neural networks (SNNs) offer biologically inspired event-driven processing with potential energy benefits, their practical advantage over artificial neural networks (ANNs) for VSP tasks remains unproven. In this work, we introduce a novel framework, ViStream, which leverages the Law of Charge Conservation (LoCC) property in ST-BIF neurons and a differential encoding (DiffEncode) scheme to optimize SNN inference for VSP. By encoding temporal differences between neighboring frames and eliminating frequent membrane resets, ViStream achieves significant computational reduction while maintaining accuracy equivalent to its ANN counterpart. We provide theoretical proofs of equivalence and validate ViStream across diverse VSP tasks, including object detection, tracking, and segmentation, demonstrating substantial energy savings without compromising performance. ViStream is publicly available at: https://github.com/Intelligent-Computing-Research-Group/ViStream
Kang You, Ziling Wei, Qinghai Guo, Zhezhi He
CVPR1
2025 BiNeuroRAM: Energy-Efficient ReRAM-Based PIM for Accurate Bipolar Spiking Neural Network Acceleration
abstract
ReRAM is a promising non-volatile memory for neuromor-phic accelerators, yet it faces challenges such as high sensing power and accuracy degradation. This work proposes BiNeuroRAM, a novel spiking neural network (SNN) accelerator leveraging ReRAM-based processing-in-memory (PIM), with three key contributions: (1) It is the first to support higher-accuracy spike-tracing bipolar-integrate-and-fire (ST-BIF) neurons, achieving 80.9% accuracy on ImageNet, 8.4% higher than the previous state-of-the-art; (2) It introduces a low-power voltage sense amplifier (LPVSA) that reduces ReRAM read power by 14.7~58.2×, enhancing energy efficiency; (3) It employs an asynchronous micro-architecture that fully exploits the event-driven nature of SNNs. Experimental results show that BiNeuroRAM improves throughput density and energy efficiency by 2.08× and 2.09× on ImageNet with ResNet-18, compared to traditional integrate-and-fire (IF) neuron-based SNN accelerators.
Jun Yan Lee, Chen Nie, Kang You, Yueyang Jia, Zhezhi He
DAC3
2025 Efficient LiDAR Reflectance Compression via Scanning Serialization
abstract
Reflectance attributes in LiDAR point clouds provide essential information for downstream tasks but remain underexplored in neural compression methods. To address this, we introduce SerLiC, a serialization-based neural compression framework to fully exploit the intrinsic characteristics of LiDAR reflectance. SerLiC first transforms 3D LiDAR point clouds into 1D sequences via scan-order serialization, offering a device-centric perspective for reflectance analysis. Each point is then tokenized into a contextual representation comprising its sensor scanning index, radial distance, and prior reflectance, for effective dependencies exploration. For efficient sequential modeling, Mamba is incorporated with a dual parallelization scheme, enabling simultaneous autoregressive dependency capture and fast processing. Extensive experiments demonstrate that SerLiC attains over 2$\times$ volume reduction against the original reflectance data, outperforming the state-of-the-art method by up to 22% reduction of compressed bits while using only 2% of its parameters. Moreover, a lightweight version of SerLiC achieves $\geq 10$ fps (frames per second) with just 111K parameters, which is attractive for real applications.
Kang You, Dandan Ding, Zhan Ma 0001
ICML2
2025 PolymorPIC: Embedding Polymorphic Processing-in-Cache in RISC-V based Processor for Full-stack Efficient AI Inference
Ziling Wei, Jun Yan Lee, Chen Nie, Kang You, Zhezhi He
MICRO5
2025 ConPCAC: Conditional Lossless Point Cloud Attribute Compression via Spatial Decomposition
abstract
A conditional lossless point cloud attribute compression method, dubbed ConPCAC, is proposed. The previous work typically codes point attributes in a point cloud in an autoregressive way, incurring unbearable coding time. By contrast, ConPCAC proposes a group-wise conditional entropy model for fast coding while preserving coding performance. Specifically, ConPCAC adopts a “Group Decomposition - Attribute Initialization - Latent Distribution Prediction” framework. First, it flexibly decomposes the original point cloud into multiple groups according to the geometry coordinate distribution. Then, the first group is coded using a base coder, e.g., the standardized G-PCC, and the following groups are progressively coded using a neural coder conditioned on their preceding groups. Two key units, Attribute Initialization (Init) and Latent Distribution Prediction (LDP), are devised in the neural coder. The Init unit employs the nearest neighbor to initialize the attributes of a group, and the LDP unit further predicts the attribute probability distribution for the group. In this way, ConPCAC enables full correlation exploration across groups and parallel processing among points in a group. Finally, the predicted probabilities are fed into the arithmetic engine to code the true attribute values of each group. Extensive experiments demonstrate the performance of ConPCAC. It achieves 14.59%, 10.32%, and 12.26% improvements over the latest G-PCC on the widely used 8iVFB, Owlii, and MVUB datasets, respectively, significantly outperforming state-of-the-art lossless PCAC methods. Moreover, its computational complexity is comparable to G-PCC and much lower than existing learning-based methods. Associated code and models will be released on the websitehttps://github.com/3dpcc/ConPCAC.
Tong Chen 0004, Kang You, Dandan Ding, Zhan Ma 0001
IEEE Trans. Circuits Syst. Video Technol.3
2024 BKDSNN: Enhancing the Performance of Learning-Based Spiking Neural Networks Training with Blurred Knowledge Distillation
Kang You, Qinghai Guo, Zhezhi He
ECCV (50)2
2024 Obtaining Optimal Spiking Neural Network in Sequence Learning via CRNN-SNN Conversion
Jiahao Su, Kang You, Weizhi Xu 0001, Zhezhi He
ICANN (10)2
2024 SpikeZIP-TF: Conversion is All You Need for Transformer-based SNN
abstract
Spiking neural network (SNN) has attracted great attention due to its characteristic of high efficiency and accuracy. Currently, the ANN-to-SNN conversion methods can obtain ANN on-par accuracy SNN with ultra-low latency (8 time-steps) in CNN structure on computer vision (CV) tasks. However, as Transformer-based networks have achieved prevailing precision on both CV and natural language processing (NLP), the Transformer-based SNNs are still encounting the lower accuracy w.r.t the ANN counterparts. In this work, we introduce a novel ANN-to-SNN conversion method called SpikeZIP-TF, where ANN and SNN are exactly equivalent, thus incurring no accuracy degradation. SpikeZIP-TF achieves 83.82% accuracy on CV dataset (ImageNet) and 93.79% accuracy on NLP dataset (SST-2), which are higher than SOTA Transformer-based SNNs. The code is available in GitHub: https://github.com/Intelligent-Computing-Research-Group/SpikeZIP_transformer
Kang You, Chen Nie, Zhijie Deng, Qinghai Guo, Zhezhi He
ICML1
2024 Pointsoup: High-Performance and Extremely Low-Decoding-Latency Learned Geometry Codec for Large-Scale Point Cloud Scenes
Kang You, Li Yu 0004, Pan Gao 0001, Dandan Ding
IJCAI1
2023 Raw Ultrasound-Based Phonetic Segments Classification Via Mask Modeling
abstract
Ultrasound tongue imaging is widely used in clinical linguistics and phonetics. Recently, deep neural networks, especially convolutional neural networks, have been widely used in the interpretation and analysis of ultrasound tongue images (UTI). Despite achieving satisfactory performance, deep models rely on a large amount of manually labeled data, which is often difficult to obtain in practical settings. To address this issue, this paper focuses on how to utilize a large amount of unlabeled UTI data to improve the performance of UTI classification task. Specifically, we explore self-supervised learning with masking modeling strategy. By predicting the masked part, our pre-trained model enables the neural network to infer contextual information. Then, we fine-tune the pre-trained model with a small amount of labeled data. Compared with the previous competing algorithms, our method can improve the classification accuracy by an average of 13.33% in four different scenarios.
Kang You, Bo Liu 0014, Kele Xu, Yunsheng Xiong, Qisheng Xu, Ming Feng, Tamás Gábor Csapó, Boqing Zhu
ICASSP1
2022 Masked Modeling-based Audio Representation for ACM Multimedia 2022 Computational Paralinguistics ChallengE
abstract
In this paper, we present our solution for ACM Multimedia 2022 Computational Paralinguistics Challenge. Our method employs the self-supervised learning paradigm, as it achieves promising results in computer vision and audio signal processing. Specifically, we firstly explore modifying the Swin Transformer architecture to learn general representation for the audio signals, accompanied with random masking on the log-mel spectrogram. The main goal of the pretext task is to predict the masked parts, by combining the advantages of the Swin-Transformer and masked modeling. For the downstream tasks, we utilize the labelled datasets to fine-tune the pre-trained model. Compared with the competitive baselines, our approach can provide significant performance improvements without ensembling.
Kang You, Kele Xu, Boqing Zhu, Ming Feng, Bo Liu 0014, Bo Ding 0001
ACM Multimedia1
2021 Patch-Based Deep Autoencoder for Point Cloud Geometry Compression
abstract
The ever-increasing 3D application makes the point cloud compression unprecedentedly important and needed. In this paper, we propose a patch-based compression process using deep learning, focusing on the lossy point cloud geometry compression. Unlike existing point cloud compression networks, which apply feature extraction and reconstruction on the entire point cloud, we divide the point cloud into patches and compress each patch independently. In the decoding process, we finally assemble the decompressed patches into a complete point cloud. In addition, we train our network by a patch-to-patch criterion, i.e., use the local reconstruction loss for optimization, to approximate the global reconstruction optimality. Our method outperforms the state-of-the-art in terms of rate-distortion performance, especially at low bitrates. Moreover, the compression process we proposed can guarantee to generate the same number of points as the input. The network model of this method can be easily applied to other point cloud reconstruction problems, such as upsampling.
Kang You, Pan Gao 0001
MMAsia1