EDBT 2026 Demo / reviewers in the wild / expert
Yuhan Zhang 0006
dblp:06/7406-6
· DBLP profile ↗
15ranked-venue papers
4as first author
15since 2021 · last 2026
0000-0001-8579-4943ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 1 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Generative Model for 2.5D-Assisted Future Urban Remote Sensing Image SynthesisabstractGenerating realistic future urban remote sensing imagery is critical for visualizing potential urban changes and supporting related technical analysis within urban planning. Traditional 2D-assisted methods are inherently limited in synthesize vertical development and infrastructure evolution, as they rely on binary planning maps. To address these limitations, we propose a novel 2.5D-assisted future urban remote sensing image synthesis method, aimed at generating future urban layouts based on existing urban structures and 2.5D planning maps. Specifically, the 2.5D map is divided into construction and demolition components, which are then integrated with the existing layout images and the embedding of the corresponding text as conditions for our generative model. We further design two trainable cascaded gated attention layers that process these two conditions separately and embed them into the latent diffusion model (LDM). This approach allows our model to dynamically comprehend the planning design requirements for key areas, making adjustments to accommodate diverse demands. Compared to existing state-of-the-art (SoTA) methods, our approach effectively targets design requirements, enabling flexible modifications that involve new constructions and demolitions in relevant urban areas. Experimental results on the 3DCD dataset demonstrate that the images generated by our method retain high fidelity and exhibit strong consistency with the 2.5D planning map. Yuhan Zhang 0006, Jie Zhou 0001, Weihang Peng 0001, Xiaode Liu, Yuanpei Chen |
IEEE Signal Process. Lett. | 1 |
| 2025 | Spiking Transformer: Introducing Accurate Addition-Only Spiking Self-Attention for TransformerabstractTransformers have demonstrated outstanding performance across a wide range of tasks, owing to their self-attention mechanism, but they are highly energy-consuming. Spiking Neural Networks have emerged as a promising energy-efficient alternative to traditional Artificial Neural Networks, leveraging event-driven computation and binary spikes for information transfer. The combination of Transformers’ capabilities with the energy efficiency of SNNs offers a compelling opportunity. This paper addresses the challenge of adapting the self-attention mechanism of Transformers to the spiking paradigm by introducing a novel approach: Accurate Addition-Only Spiking Self-Attention (A2OS2A). Unlike existing methods that rely solely on binary spiking neurons for all components of the self-attention mechanism, our approach integrates binary, ReLU, and ternary spiking neurons. This hybrid strategy significantly improves accuracy while preserving non-multiplicative computations. Moreover, our method eliminates the need for softmax and scaling operations. Extensive experiments show that the A2OS2A-based Spiking Transformer outperforms existing SNN-based Transformers on several datasets, even achieving an accuracy of 78.66% on ImageNet-1K. Our work represents a significant advancement in SNN-based Transformer models, offering a more accurate and efficient solution for real-world applications. Yufei Guo 0001, Xiaode Liu, Yuanpei Chen, Weihang Peng 0001, Yuhan Zhang 0006, Zhe Ma 0001 |
CVPR | 5 |
| 2025 | ReverB-SNN: Reversing Bit of the Weight and Activation for Spiking Neural NetworksabstractThe Spiking Neural Network (SNN), a biologically inspired neural network infrastructure, has garnered significant attention recently. SNNs utilize binary spike activations for efficient information transmission, replacing multiplications with additions, thereby enhancing energy efficiency. However, binary spike activation maps often fail to capture sufficient data information, resulting in reduced accuracy.
To address this challenge, we advocate reversing the bit of the weight and activation, called \textbf{ReverB}, inspired by recent findings that highlight greater accuracy degradation from quantizing activations compared to weights. Specifically, our method employs real-valued spike activations alongside binary weights in SNNs. This preserves the event-driven and multiplication-free advantages of standard SNNs while enhancing the information capacity of activations.
Additionally, we introduce a trainable factor within binary weights to adaptively learn suitable weight amplitudes during training, thereby increasing network capacity. To maintain efficiency akin to vanilla \textbf{ReverB}, our trainable binary weight SNNs are converted back to standard form using a re-parameterization technique during inference.
Extensive experiments across various network architectures and datasets, both static and dynamic, demonstrate that our approach consistently outperforms state-of-the-art methods. Yufei Guo 0001, Yuhan Zhang 0006, Jie Zhou 0001, Xiaode Liu, Yuanpei Chen, Weihang Peng 0001, Zhe Ma 0001 |
ICML | 2 |
| 2025 | UP-Diff: Latent Diffusion Model for Remote Sensing Urban PredictionabstractRemote sensing (RS) technology has become essential for monitoring urban development, including applications like population growth analysis, transportation congestion forecasting, and climate change detection (CD). However, its potential for future urban planning (UP), particularly in predicting future urban layouts remains largely unexplored. This study introduces UP-Diff, a novel method leveraging generative models for UP, to address this gap. UP-Diff leverages information from current urban layouts and planned change maps to predict future urban configurations. Key challenges addressed include the integration of urban layouts and change maps into latent diffusion model (LDM) through careful architecture improvements and the mitigation of limited training data by employing a pretrained stable diffusion (SD) model with fixed weights, trainable ConvNeXt, and trainable cross-attention layers. Our method significantly streamlines the UP process by automating layout predictions, thus reducing the time and effort required compared to traditional manual methods. Comprehensive evaluations on the learning, vision, and RS dataset (LEVIR-CD) and Sun Yat-Sen University dataset (SYSU-CD) validate that UP-Diff achieves high-fidelity predictions of future urban layouts, demonstrating its effectiveness and potential for advancing RS-based UP methodologies. Our code and model weights are available athttps://github.com/zeyuwang-zju/UP-Diff. Zeyu Wang 0010, Zecheng Hao, Yuhan Zhang 0006, Yuchao Feng, Yufei Guo 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2025 | PT-BitNet: Scaling up the 1-Bit large language model with post-training quantization
Yufei Guo 0001, Zecheng Hao, Jiahang Shao, Jie Zhou 0001, Xiaode Liu, Yuhan Zhang 0006, Yuanpei Chen, Weihang Peng 0001, Zhe Ma 0001 |
Neural Networks | 7 |
| 2024 | Ternary Spike: Learning Ternary Spikes for Spiking Neural NetworksabstractThe Spiking Neural Network (SNN), as one of the biologically inspired neural network infrastructures, has drawn increasing attention recently. It adopts binary spike activations to transmit information, thus the multiplications of activations and weights can be substituted by additions, which brings high energy efficiency. However, in the paper, we theoretically and experimentally prove that the binary spike activation map cannot carry enough information, thus causing information loss and resulting in accuracy decreasing. To handle the problem, we propose a ternary spike neuron to transmit information. The ternary spike neuron can also enjoy the event-driven and multiplication-free operation advantages of the binary spike neuron but will boost the information capacity. Furthermore, we also embed a trainable factor in the ternary spike neuron to learn the suitable spike amplitude, thus our SNN will adopt different spike amplitudes along layers, which can better suit the phenomenon that the membrane potential distributions are different along layers. To retain the efficiency of the vanilla ternary spike, the trainable ternary spike SNN will be converted to a standard one again via a re-parameterization technique in the inference. Extensive experiments with several popular network structures over static and dynamic datasets show that the ternary spike can consistently outperform state-of-the-art methods. Our code is open-sourced at https://github.com/yfguo91/Ternary-Spike. Yufei Guo 0001, Yuanpei Chen, Xiaode Liu, Weihang Peng 0001, Yuhan Zhang 0006, Xuhui Huang, Zhe Ma 0001 |
AAAI | 5 |
| 2024 | Enhancing Representation of Spiking Neural Networks via Similarity-Sensitive Contrastive LearningabstractSpiking neural networks (SNNs) have attracted intensive attention as a promising energy-efficient alternative to conventional artificial neural networks (ANNs) recently, which could transmit information in form of binary spikes rather than continuous activations thus the multiplication of activation and weight could be replaced by addition to save energy. However, the binary spike representation form will sacrifice the expression performance of SNNs and lead to accuracy degradation compared with ANNs. Considering improving feature representation is beneficial to training an accurate SNN model, this paper focuses on enhancing the feature representation of the SNN. To this end, we establish a similarity-sensitive contrastive learning framework, where SNN could capture significantly more information from its ANN counterpart to improve representation by Mutual Information (MI) maximization with layer-wise sensitivity to similarity. In specific, it enriches the SNN’s feature representation by pulling the positive pairs of SNN's and ANN's feature representation of each layer from the same input samples closer together while pushing the negative pairs from different samples further apart. Experimental results show that our method consistently outperforms the current state-of-the-art algorithms on both popular non-spiking static and neuromorphic datasets. Yuhan Zhang 0006, Xiaode Liu, Yuanpei Chen, Weihang Peng 0001, Yufei Guo 0001, Xuhui Huang, Zhe Ma 0001 |
AAAI | 1 |
| 2024 | EnOF-SNN: Training Accurate Spiking Neural Networks via Enhancing the Output FeatureabstractSpiking neural networks (SNNs) have gained more and more interest as one of the energy-efficient alternatives of conventional artificial neural networks (ANNs). They exchange 0/1 spikes for processing information, thus most of the multiplications in networks can be replaced by additions. However, binary spike feature maps will limit the expressiveness of the SNN and result in unsatisfactory performance compared with ANNs.
It is shown that a rich output feature representation, i.e., the feature vector before classifier) is beneficial to training an accurate model in ANNs for classification.
We wonder if it also does for SNNs and how to improve the feature representation of the SNN.
To this end, we materialize this idea in two special designed methods for SNNs.
First, inspired by some ANN-SNN methods that directly copy-paste the weight parameters from trained ANN with light modification to homogeneous SNN can obtain a well-performed SNN, we use rich information of the weight parameters from the trained ANN counterpart to guide the feature representation learning of the SNN.
In particular, we present the SNN's and ANN's feature representation from the same input to ANN's classifier to product SNN's and ANN's outputs respectively and then align the feature with the KL-divergence loss as in knowledge distillation methods, called L_ AF loss.
It can be seen as a novel and effective knowledge distillation method specially designed for the SNN that comes from both the knowledge distillation and ANN-SNN methods. Various ablation study shows that the L_AF loss is more powerful than the vanilla knowledge distillation method.
Second, we replace the last Leaky Integrate-and-Fire (LIF) activation layer as the ReLU activation layer to generate the output feature, thus a more powerful SNN with full-precision feature representation can be achieved but with only a little extra computation.
Experimental results show that our method consistently outperforms the current state-of-the-art algorithms on both popular non-spiking static and neuromorphic datasets. We provide an extremely simple but effective way to train high-accuracy spiking neural networks. Yufei Guo 0001, Weihang Peng 0001, Xiaode Liu, Yuanpei Chen, Yuhan Zhang 0006, Zhou Jie, Zhe Ma 0001 |
NeurIPS | 5 |
| 2024 | Take A Shortcut Back: Mitigating the Gradient Vanishing for Training Spiking Neural NetworksabstractThe Spiking Neural Network (SNN) is a biologically inspired neural network infrastructure that has recently garnered significant attention. It utilizes binary spike activations to transmit information, thereby replacing multiplications with additions and resulting in high energy efficiency. However, training an SNN directly poses a challenge due to the undefined gradient of the firing spike process. Although prior works have employed various surrogate gradient training methods that use an alternative function to replace the firing process during back-propagation, these approaches ignore an intrinsic problem: gradient vanishing. To address this issue, we propose a shortcut back-propagation method in the paper, which advocates for transmitting the gradient directly from the loss to the shallow layers. This enables us to present the gradient to the shallow layers directly, thereby significantly mitigating the gradient vanishing problem. Additionally, this method does not introduce any burden during the inference phase.
To strike a balance between final accuracy and ease of training, we also propose an evolutionary training framework and implement it by inducing a balance coefficient that dynamically changes with the training epoch, which further improves the network's performance. Extensive experiments conducted over static and dynamic datasets using several popular network structures reveal that our method consistently outperforms state-of-the-art methods. Yufei Guo 0001, Yuanpei Chen, Zecheng Hao, Weihang Peng 0001, Zhou Jie, Yuhan Zhang 0006, Xiaode Liu, Zhe Ma 0001 |
NeurIPS | 6 |
| 2023 | Audio-Driven High Definetion and Lip-Synchronized Talking Face Generation Based on Face ReenactmentabstractGenerating audio-driven photo-realistic talking face has received intensive attention due to its ability to bring more new human-computer interaction experiences. However, previous works struggled to balance high definition, lip synchronization, and low customization costs, which would degrade the user experience. In this paper, a novel audio-driven talking face generation method was proposed, which subtly converts the problem of improving video definition into the problem of face reenactment to produce both lip-synchronized and high- definition face video. The framework is decoupled, meaning that the same trained model can be used on arbitrary characters and audio without further customizing training for specific people, thus significantly reducing costs. Experiment results show that our proposed method achieves the high video definition, and comparable lip synchronization performance with the existing state-of-the-art methods. Yuhan Zhang 0006, Weihua He, Yaoyuan Wang, Shunbo Zhou |
ICASSP | 2 |
| 2023 | RMP-Loss: Regularizing Membrane Potential Distribution for Spiking Neural NetworksabstractSpiking Neural Networks (SNNs) as one of the biology-inspired models have received much attention recently. It can significantly reduce energy consumption since they quantize the real-valued membrane potentials to 0/1 spikes to transmit information thus the multiplications of activations and weights can be replaced by additions when implemented on hardware. However, this quantization mechanism will inevitably introduce quantization error, thus causing catastrophic information loss. To address the quantization error problem, we propose a regularizing membrane potential loss (RMP-Loss) to adjust the distribution which is directly related to quantization error to a range close to the spikes. Our method is extremely simple to implement and straightforward to train an SNN. Furthermore, it is shown to consistently outperform previous state-of-the-art methods over different network architectures and datasets. Yufei Guo 0001, Xiaode Liu, Yuanpei Chen, Liwen Zhang 0001, Weihang Peng 0001, Yuhan Zhang 0006, Xuhui Huang, Zhe Ma 0001 |
ICCV | 6 |
| 2023 | Membrane Potential Batch Normalization for Spiking Neural NetworksabstractAs one of the energy-efficient alternatives of conventional neural networks (CNNs), spiking neural networks (SNNs) have gained more and more interest recently. To train the deep models, some effective batch normalization (BN) techniques are proposed in SNNs. All these BNs are suggested to be used after the convolution layer as usually doing in CNNs. However, the spiking neuron is much more complex with the spatio-temporal dynamics. The regulated data flow after the BN layer will be disturbed again by the membrane potential updating operation before the firing function, i.e., the nonlinear activation. Therefore, we advocate adding another BN layer before the firing function to normalize the membrane potential again, called MPBN. To eliminate the induced time cost of MPBN, we also propose a training-inference-decoupled re-parameterization technique to fold the trained MPBN into the firing threshold. With the re-parameterization technique, the MPBN will not introduce any extra time burden in the inference. Furthermore, the MPBN can also adopt the element-wised form, while these BNs after the convolution layer can only use the channel-wised form. Experimental results show that the proposed MPBN performs well on both popular non-spiking static and neuromorphic datasets. Yufei Guo 0001, Yuhan Zhang 0006, Yuanpei Chen, Weihang Peng 0001, Xiaode Liu, Liwen Zhang 0001, Xuhui Huang, Zhe Ma 0001 |
ICCV | 2 |
| 2023 | Spiking PointNet: Spiking Neural Networks for Point CloudsabstractRecently, Spiking Neural Networks (SNNs), enjoying extreme energy efficiency, have drawn much research attention on 2D visual recognition and shown gradually increasing application potential. However, it still remains underexplored whether SNNs can be generalized to 3D recognition. To this end, we present Spiking PointNet in the paper, the first spiking neural model for efficient deep learning on point clouds. We discover that the two huge obstacles limiting the application of SNNs in point clouds are: the intrinsic optimization obstacle of SNNs that impedes the training of a big spiking model with large time steps, and the expensive memory and computation cost of PointNet that makes training a big spiking point model unrealistic. To solve the problems simultaneously, we present a trained-less but learning-more paradigm for Spiking PointNet with theoretical justifications and in-depth experimental analysis. In specific, our Spiking PointNet is trained with only a single time step but can obtain better performance with multiple time steps inference, compared to the one trained directly with multiple time steps. We conduct various experiments on ModelNet10, ModelNet40 to demonstrate the effectiveness of Sipiking PointNet. Notably, our Spiking PointNet even can outperform its ANN counterpart, which is rare in the SNN field thus providing a potential research direction for the following work. Moreover, Spiking PointNet shows impressive speedup and storage saving in the training phase. Our code is open-sourced at https://github.com/DayongRen/Spiking-PointNet. Dayong Ren, Zhe Ma 0001, Yuanpei Chen, Weihang Peng 0001, Xiaode Liu, Yuhan Zhang 0006, Yufei Guo 0001 |
NeurIPS | 6 |
| 2022 | Meta Talk: Learning To Data-Efficiently Generate Audio-Driven Lip-Synchronized Talking Face With High DefinitionabstractAudio-driven talking face, driving talking face by audio, has received considerable attention in multi-modal learning due to its widespread use in virtual reality. However, long-time recording of target high-quality video is needed by most existing audio-driven talking face studies, which significantly increases customization costs. This paper proposes a novel data-efficient audio-driven talking face generation method, which uses just a short target video to produce both lip-synchronized and high-definition face video driven by arbitrary audio in the wild. Current methods suffer from many problems, such as low definition, asynchronization of lip movement and voice, and intense demands for videos for training. In this work, the original target character’s face images are decomposed into 3D face model parameters including expression, geometry, illumination, etc. Then, low-definition pseudo video generated by an adapted target face video bridges the powerful pre-trained audio-driven model to our audio-to-expression transformation network and help to transfer the ability of audio-identity disentanglement. The expression is replaced via an audio and then combined with other face parameters to render a synthetic face. Finally, a neural rendering network translates the synthetic face into talking face without loss of definition. Experimental results show that the proposed method has the best performance in high-definition image quality, and comparable performance in lip synchronization compared with the existing state-of-the-art methods. Yuhan Zhang 0006, Weihua He, Yaoyuan Wang, Jianxing Liao |
ICASSP | 1 |
| 2022 | A Data-Driven Modeling Method for Stochastic Nonlinear Degradation Process With Application to RUL EstimationabstractThis article proposes a novel modeling method for the stochastic nonlinear degradation process by using the relevance vector machine (RVM), which can describe the nonlinearity of degradation process more flexibly and accurately. Compared with the existing methods, where degradation processes are modeled as the Wiener process with a nonlinear drift function formulized as the power law or exponential law, this kind of modeling method can characterize degradation processes with more nonlinear behavior. Instead of modeling the drift coefficient of the Wiener process directly, the weighted combination of basis functions is utilized to express the increment of the Wiener process and the parameters are calculated by a sparse Bayesian learning algorithm. Based on the proposed model, a numerical approximation formula for the probability density function (PDF) of the remaining useful life (RUL) is derived. Finally, comparison studies, including a numerical simulation and a practical case, are provided to demonstrate the effectiveness and the accuracy of the proposed methods for RUL estimation. Yuhan Zhang 0006, Ying Yang 0002, He Li 0025, Xianchao Xiu, Wanquan Liu |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |