Liwen Zhang 0001

dblp:94/905-1 · DBLP profile ↗
← Back
17ranked-venue papers
6as first author
12since 2021 · last 2026
0000-0001-8457-2943ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 8 since 2021
YearPublicationVenuePosition
2026 AXFL: Axial prior-guided cross-view fusion learning for radar semantic segmentation
Liwen Zhang 0001, Youcheng Zhang, Qingmin Liao
Expert Syst. Appl.2
2025 InfoMin-based Query Embedding Optimization For Query-based Universal Sound Separation
abstract
The query-based universal sound separation (QUSS) has been addressed, aiming to perform the separation of specific sound sources based on a given query. Most of existed methods focus on the improvement of separation models, ignoring the influence of category-conditioned query embedding distribution on separation performance. To address this issue, we propose an optimization method for query embedding that reduces mutual information (MI) between query embeddings while keeping task-related information intact, named the InfoMin principle. In addition, we propose the Frequency-varying Feature-wise Linear Modulation (FFiLM), which leverages frequency band differences in acoustic events to enhance the modulation capability of query embedding and improve the performance of the separation model. Experimental results show that our method achieves considerable improvements over the existing SoTA method.
Jiqing Han 0001, Liwen Zhang 0001, Youcheng Zhang
ICASSP3
2024 SpecAR-Net: Spectrogram Analysis and Representation Network for Time Series
Liwen Zhang 0001, Youcheng Zhang, Shi Peng, Zhe Ma 0001
IJCAI2
2024 TARSS-Net: Temporal-Aware Radar Semantic Segmentation Network
abstract
Radar signal interpretation plays a crucial role in remote detection and ranging. With the gradual display of the advantages of neural network technology in signal processing, learning-based radar signal interpretation is becoming a research hot-spot and made great progress. And since radar semantic segmentation (RSS) can provide more fine-grained target information, it has become a more concerned direction in this field. However, the temporal information, which is an important clue for analyzing radar data, has not been exploited sufficiently in present RSS frameworks. In this work, we propose a novel temporal information learning paradigm, i.e., data-driven temporal information aggregation with learned target-history relations. Following this idea, a flexible learning module, called Temporal Relation-Aware Module (TRAM) is carefully designed. TRAM contains two main blocks: i) an encoder for capturing the target-history temporal relations (TH-TRE) and ii) a learnable temporal relation attentive pooling (TRAP) for aggregating temporal information. Based on TRAM, an end-to-end Temporal-Aware RSS Network (TARSS-Net) is presented, which has outstanding performance on publicly available and our collected real-measured datasets. Code and supplementary materials are available at https://github.com/zlw9161/TARSS-Net.
Youcheng Zhang, Liwen Zhang 0001, ZijunHu, Pengcheng Pi, Yuanpei Chen, Shi Peng, Zhe Ma 0001
NeurIPS2
2023 PeakConv: Learning Peak Receptive Field for Radar Semantic Segmentation
abstract
The modern machine learning-based technologies have shown considerable potential in automatic radar scene understanding. Among these efforts, radar semantic segmentation (RSS) can provide more refined and detailed information including the moving objects and background clutters within the effective receptive field of the radar. Motivated by the success of convolutional networks in various visual computing tasks, these networks have also been introduced to solve RSS task. However, neither the regular convolution operation nor the modified ones are specific to interpret radar signals. The receptive fields of existing convolutions are defined by the object presentation in optical signals, but these two signals have different perception mechanisms. In classic radar signal processing, the object signature is detected according to a local peak response, i.e., CFAR detection. Inspired by this idea, we redefine the receptive field of the convolution operation as the peak receptive field (PRF) and propose the peak convolution operation (PeakConv) to learn the object signatures in an end-to-end network. By incorporating the proposed PeakConv layers into the encoders, our RSS network can achieve better segmentation results compared with other SoTA methods on a multi-view real-measured dataset collected from an FMCW radar. Our code for PeakConv is available at https://github.com/zlw9161/PKC.
Liwen Zhang 0001, Youcheng Zhang, Yufei Guo 0001, Yuanpei Chen, Xuhui Huang, Zhe Ma 0001
CVPR1
2023 RMP-Loss: Regularizing Membrane Potential Distribution for Spiking Neural Networks
abstract
Spiking Neural Networks (SNNs) as one of the biology-inspired models have received much attention recently. It can significantly reduce energy consumption since they quantize the real-valued membrane potentials to 0/1 spikes to transmit information thus the multiplications of activations and weights can be replaced by additions when implemented on hardware. However, this quantization mechanism will inevitably introduce quantization error, thus causing catastrophic information loss. To address the quantization error problem, we propose a regularizing membrane potential loss (RMP-Loss) to adjust the distribution which is directly related to quantization error to a range close to the spikes. Our method is extremely simple to implement and straightforward to train an SNN. Furthermore, it is shown to consistently outperform previous state-of-the-art methods over different network architectures and datasets.
Yufei Guo 0001, Xiaode Liu, Yuanpei Chen, Liwen Zhang 0001, Weihang Peng 0001, Yuhan Zhang 0006, Xuhui Huang, Zhe Ma 0001
ICCV4
2023 Membrane Potential Batch Normalization for Spiking Neural Networks
abstract
As one of the energy-efficient alternatives of conventional neural networks (CNNs), spiking neural networks (SNNs) have gained more and more interest recently. To train the deep models, some effective batch normalization (BN) techniques are proposed in SNNs. All these BNs are suggested to be used after the convolution layer as usually doing in CNNs. However, the spiking neuron is much more complex with the spatio-temporal dynamics. The regulated data flow after the BN layer will be disturbed again by the membrane potential updating operation before the firing function, i.e., the nonlinear activation. Therefore, we advocate adding another BN layer before the firing function to normalize the membrane potential again, called MPBN. To eliminate the induced time cost of MPBN, we also propose a training-inference-decoupled re-parameterization technique to fold the trained MPBN into the firing threshold. With the re-parameterization technique, the MPBN will not introduce any extra time burden in the inference. Furthermore, the MPBN can also adopt the element-wised form, while these BNs after the convolution layer can only use the channel-wised form. Experimental results show that the proposed MPBN performs well on both popular non-spiking static and neuromorphic datasets.
Yufei Guo 0001, Yuhan Zhang 0006, Yuanpei Chen, Weihang Peng 0001, Xiaode Liu, Liwen Zhang 0001, Xuhui Huang, Zhe Ma 0001
ICCV6
2023 Joint A-SNN: Joint training of artificial and spiking neural networks via self-Distillation and weight factorization
Yufei Guo 0001, Weihang Peng 0001, Yuanpei Chen, Liwen Zhang 0001, Xiaode Liu, Xuhui Huang, Zhe Ma 0001
Pattern Recognit.4
2022 RecDis-SNN: Rectifying Membrane Potential Distribution for Directly Training Spiking Neural Networks
abstract
The brain-inspired and event-driven Spiking Neural Network (SNN) aiming at mimicking the synaptic activity of biological neurons has received increasing attention. It transmits binary spike signals between network units when the membrane potential exceeds the firing threshold. This biomimetic mechanism of SNN appears energy-efficiency with its power sparsity and asynchronous operations on spike events. Unfortunately, with the propagation of binary spikes, the distribution of membrane potential will shift, leading to degeneration, saturation, and gradient mismatch problems, which would be disadvantageous to the network optimization and convergence. Such undesired shifts would prevent the SNN from performing well and going deep. To tackle these problems, we attempt to rectify the membrane potential distribution (MPD) by designing a novel distribution loss, MPD-Loss, which can explicitly penalize the un-desired shifts without introducing any additional operations in the inference phase. Moreover, the proposed method can also mitigate the quantization error in SNNs, which is usually ignored in other works. Experimental results demonstrate that the proposed method can directly train a deeper, larger, and better-performing SNN within fewer timesteps.
Yufei Guo 0001, Xinyi Tong 0001, Yuanpei Chen, Liwen Zhang 0001, Xiaode Liu, Zhe Ma 0001, Xuhui Huang
CVPR4
2022 Reducing Information Loss for Spiking Neural Networks
Yufei Guo 0001, Yuanpei Chen, Liwen Zhang 0001, YingLei Wang, Xiaode Liu, Xinyi Tong 0001, Yuanyuan Ou, Xuhui Huang, Zhe Ma 0001
ECCV (11)3
2022 Real Spike: Learning Real-Valued Spikes for Spiking Neural Networks
Yufei Guo 0001, Liwen Zhang 0001, Yuanpei Chen, Xinyi Tong 0001, Xiaode Liu, YingLei Wang, Xuhui Huang, Zhe Ma 0001
ECCV (12)2
2022 IM-Loss: Information Maximization Loss for Spiking Neural Networks
abstract
Spiking Neural Network (SNN), recognized as a type of biologically plausible architecture, has recently drawn much research attention. It transmits information by $0/1$ spikes. This bio-mimetic mechanism of SNN demonstrates extreme energy efficiency since it avoids any multiplications on neuromorphic hardware. However, the forward-passing $0/1$ spike quantization will cause information loss and accuracy degradation. To deal with this problem, the Information maximization loss (IM-Loss) that aims at maximizing the information flow in the SNN is proposed in the paper. The IM-Loss not only enhances the information expressiveness of an SNN directly but also plays a part of the role of normalization without introducing any additional operations (\textit{e.g.}, bias and scaling) in the inference phase. Additionally, we introduce a novel differentiable spike activity estimation, Evolutionary Surrogate Gradients (ESG) in SNNs. By appointing automatic evolvable surrogate gradients for spike activity function, ESG can ensure sufficient model updates at the beginning and accurate gradients at the end of the training, resulting in both easy convergence and high task performance. Experimental results on both popular non-spiking static and neuromorphic datasets show that the SNN models trained by our method outperform the current state-of-the-art algorithms.
Yufei Guo 0001, Yuanpei Chen, Liwen Zhang 0001, Xiaode Liu, YingLei Wang, Xuhui Huang, Zhe Ma 0001
NeurIPS3
2020 ATReSN-Net: Capturing Attentive Temporal Relations in Semantic Neighborhood for Acoustic Scene Classification
Liwen Zhang 0001, Jiqing Han 0001, Ziqiang Shi
INTERSPEECH1
2020 FurcaNeXt: End-to-End Monaural Speech Separation with Dynamic Gated Dilated Temporal Convolutional Networks
Liwen Zhang 0001, Ziqiang Shi, Jiqing Han 0001, Anyan Shi, Ding Ma 0001
MMM (1)1
2020 Learning Temporal Relations from Semantic Neighbors for Acoustic Scene Classification
abstract
Convolutional networks have achieved the state-of-the-art performance on Acoustic Scene Classification (ASC). Given the Log Mel-Spectrogram of an audio sample, the network can extract useful semantic contents in a certain range receptive field by stacking local convolutional operations. However, the temporal relations between different receptive fields are not captured explicitly. In this letter, we propose an end-to-end 3D Convolutional Neural Network (CNN) for ASC, named SeNoT-Net, which can generate effective audio representations by capturing temporal relations from semantic neighbors of different receptive fields over time. The SeNoT-Net treats the Log-Mel spectrogram as an ordered segment-level sequence. For each segment, the residual block can produce the semantic feature maps, then the semantic neighbors over time (SeNoT) module is applied to capture the relations between each feature point in the feature maps and its top-k semantic neighbors. The proposed SeNoT-Net outperforms most of the state-of-the-art CNN models on both DCASE 2018 and 2019 ASC datasets.
Liwen Zhang 0001, Jiqing Han 0001, Ziqiang Shi
IEEE Signal Process. Lett.1
2020 Pyramidal Temporal Pooling With Discriminative Mapping for Audio Classification
abstract
Audio signals are temporally-structured data, and learning their discriminative representations containing temporal information is crucial for the audio classification. In this article, we propose an audio representation learning method with a hierarchical pyramid structure called pyramidal temporal pooling (PTP) which aims to capture the temporal information of an entire audio sample. By stacking a global temporal pooling layer on multiple local temporal pooling layers, the PTP can capture the high-level temporal dynamics of the input feature sequence in an unsupervised way. Furthermore, in the top global temporal pooling layer, we jointly optimize a learnable discriminative mapping (DM) and a softmax classifier. Such that, a joint learning method for the discriminative audio representations and the classifier called DM-PTP is also presented. By treating the temporal encoding as a low-level constraint of a bi-level optimization problem, the DM-PTP can produce the discriminative representation while maintaining the temporal information of the whole sequence. For an audio sample with an arbitrary time duration, both our PTP and DM-PTP can encode the input feature sequence with arbitrary length into a fixed-length representation. Without using any data augmentation and ensemble learning methods, both PTP and DM-PTP outperform the state-of-the-art CNNs on the audio event recognition (AER) dataset, and can achieve comparable performance on the DCASE 2018 acoustic scene classification (ASC) dataset compared with other best models in the challenge.
Liwen Zhang 0001, Ziqiang Shi, Jiqing Han 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2018 Unsupervised Temporal Feature Learning Based on Sparse Coding Embedded BoAW for Acoustic Event Recognition
Liwen Zhang 0001, Jiqing Han 0001, Shiwen Deng
INTERSPEECH1