Wendong Zheng

dblp:236/3122 · DBLP profile ↗
← Back
22ranked-venue papers
12as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 7 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 3 · 3 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Shift-level regulated continual test-time adaptation framework
Hang Xu 0009, Wendong Zheng, Husheng Guo, Wenjian Wang 0001
Data Min. Knowl. Discov.3
2025 LGFormer: A Local-Global Dynamic Attention Window Transformer for Speech Emotion Recognition
abstract
Speech emotion recognition is important in intelligent human-computer interaction, but modeling to handle long-range dependencies and local emotional cues remains challenging. This paper proposes a Local-Global Dynamic Attention Window-based Transformer model (LGFormer). The Local module dynamically divides the window based on temporal significance, capturing fine-grained localized information and synthesizing features using the Global module. The model introduces a novel attention mechanism to optimize computational efficiency, making it particularly suitable for scenarios with limited computational resources. We evaluate the method on the IEMOCAP and MELD datasets, achieving 2.8% and 1.1% improvements in weighted accuracy and F1 score, respectively. Comparison experiments with multiple benchmark algorithms validate the effectiveness of the model.
Yinfeng Yu, Wendong Zheng
CSCWD4
2025 Can DBNNs Robust to Environmental Noise for Resource-constrained Scenarios?
abstract
Recently, the potential of lightweight models for resource-constrained scenarios has garnered significant attention, particularly in safety-critical tasks such as bio-electrical signal classification and B-ultrasound-assisted diagnostic. These tasks are frequently affected by environmental noise due to patient movement artifacts and inherent device noise, which pose significant challenges for lightweight models (e.g., deep binary neural networks (DBNNs)) to perform robust inference. A pertinent question arises: can a well-trained DBNN effectively resist environmental noise during inference? In this study, we find that the DBNN's robustness vulnerability comes from the binary weights and scaling factors. Drawing upon theoretical insights, we propose L1-infinite norm constraints for binary weights and scaling factors, which yield a tighter upper bound compared to existing state-of-the-art (SOTA) methods. Finally, visualization studies show that our approach introduces minimal noise perturbations at the periphery of the feature maps. Our approach outperforms the SOTA method, as validated by several experiments conducted on the bio-electrical and image classification datasets. We hope our findings can raise awareness among researchers about the environmental noise robustness of DBNNs.
Wendong Zheng, Husheng Guo
ICML1
2025 Audio-Guided Dynamic Modality Fusion with Stereo-Aware Attention for Audio-Visual Navigation
Yinfeng Yu, Fuchun Sun 0001, Wendong Zheng
ICONIP (5)5
2025 PGSTalker: Real-Time Audio-Driven Talking Head Generation via 3D Gaussian Splatting with Pixel-Aware Density Control
Tianheng Zhu, Yinfeng Yu, Fuchun Sun 0001, Wendong Zheng
ICONIP (4)5
2025 ECTSpeech: Enhancing Efficient Speech Synthesis via Easy Consistency Tuning
abstract
Diffusion models have demonstrated remarkable performance in speech synthesis, but typically require multi-step sampling, resulting in low inference efficiency. Recent studies address this issue by distilling diffusion models into consistency models, enabling efficient one-step generation. However, these approaches introduce additional training costs and rely heavily on the performance of pre-trained teacher models. In this paper, we propose ECTSpeech, a simple and effective one-step speech synthesis framework that, for the first time, incorporates the Easy Consistency Tuning (ECT) strategy into speech synthesis. By progressively tightening consistency constraints on a pre-trained diffusion model, ECTSpeech achieves high-quality one-step generation while significantly reducing training complexity. In addition, we design a multi-scale gate module (MSGate) to enhance the denoiser’s ability to fuse features at different scales. Experimental results on the LJSpeech dataset demonstrate that ECTSpeech achieves audio quality comparable to state-of-the-art methods under single-step sampling, while substantially reducing the model’s training cost and complexity.
Yinfeng Yu, Fuchun Sun 0001, Wendong Zheng
MMAsia5
2025 Iterative Residual Cross-Attention Mechanism: An Integrated Approach for Audio-Visual Navigation Tasks
abstract
Audio-visual navigation represents a significant area of research in which intelligent agents utilize egocentric visual and auditory perceptions to identify audio targets. Conventional navigation methodologies typically adopt a staged modular design, which involves first executing feature fusion, then utilizing Gated Recurrent Unit (GRU) modules for sequence modeling, and finally making decisions through reinforcement learning. While this modular approach has demonstrated effectiveness, it may also lead to redundant information processing and inconsistencies in information transmission between the various modules during the feature fusion and GRU sequence modeling phases. This paper presents IRCAM-AVN (Iterative Residual Cross-Attention Mechanism for Audiovisual Navigation), an end-to-end framework that integrates multimodal information fusion and sequence modeling within a unified IRCAM module, thereby replacing the traditional separate components for fusion and GRU. This innovative mechanism employs a multi-level residual design that concatenates initial multimodal sequences with processed information sequences. This methodological shift progressively optimizes the feature extraction process while reducing model bias and enhancing the model’s stability and generalization capabilities. Empirical results indicate that intelligent agents employing the iterative residual cross-attention mechanism exhibit superior navigation performance.
Yinfeng Yu, Fuchun Sun 0001, Wendong Zheng
SMC5
2025 EGSTalker: Real-Time Audio-Driven Talking Head Generation with Efficient Gaussian Deformation
abstract
This paper presents EGSTalker, a real-time audio-driven talking head generation framework based on 3D Gaussian Splatting (3DGS). Designed to enhance both speed and visual fidelity, EGSTalker requires only 3–5 minutes of training video to synthesize high-quality facial animations. The framework comprises two key stages: static Gaussian initialization and audio-driven deformation. In the first stage, a multi-resolution hash triplane and a Kolmogorov-Arnold Network (KAN) are used to extract spatial features and construct a compact 3D Gaussian representation. In the second stage, we propose an Efficient Spatial-Audio Attention (ESAA) module to fuse audio and spatial cues, while KAN predicts the corresponding Gaussian deformations. Extensive experiments demonstrate that EGSTalker achieves rendering quality and lip-sync accuracy comparable to state-of-the-art methods, while significantly outperforming them in inference speed. These results highlight EGSTalker’s potential for real-time multimedia applications.
Tianheng Zhu, Yinfeng Yu, Fuchun Sun 0001, Wendong Zheng
SMC5
2024 A Large-area Tactile Sensor for Distributed Force Sensing Using Highly Sensitive Piezoresistive Sponge
abstract
Tactile sensing plays a critical role in enabling robots to interact safely with target objects in dynamic and unstructured environments. While various tactile sensors based on different sensing principles or different sensitive materials have been proposed, the development of flexible large-area tactile sensors for robots is still challenging. In this paper, a novel highly sensitive piezoresistive sponge based on multi-walled carbon nanotubes (MWCNTs) and polyurethane (PU) sponge is fabricated for pressure sensing. The sensing behavior of the piezoresistive sponge was experimentally evaluated, showing high sensitivity and fast response. Based on the piezoresistive sponge, a flexible large-area tactile sensor is designed for distributed force detection with electrical resistance tomography technology. The sensing performance of the sensor is validated by touch location, sensitivity analysis, real-time touch discrimination, and touch modality recognition. The experimental results indicate that the sensor performs well in detecting the position and force of contact in a large area. The sensor’s performance shows promise in embodied tactile sensing and human–robot interaction.
Wendong Zheng, Di Guo 0002, Wuqiang Yang, Huaping Liu 0001
ICRA1
2024 Data-driven electrical resistance tomography for robotic large-area tactile sensing
Wendong Zheng, Huaping Liu 0001, Fuchun Sun 0001
Sci. China Inf. Sci.1
2023 Towards Effective Training of Robust Spiking Recurrent Neural Networks Under General Input Noise via Provable Analysis
abstract
Recently, bio-inspired spiking neural networks (SNN) with recurrent structures (SRNN) have received increasingly more attention due to their appealing properties for energy-efficiently solving time-domain classification tasks. SRNN s are often executed in noisy environments on resource-constrained devices which can however greatly compromise its accuracy. Thus, one fundamental question that remains unanswered is whether a formal analysis under the general input noise disturbances can be obtained to guarantee the robustness of SRNNs. Several studies have shown great promises by optimizing the bound over adverse perturbations based on Lipschitz continuity theorem, but most of these theoretical analysis are confined to convolutional neural networks (CNN). In this work, we take a further step towards robust SRNN training via provable robustness analysis over input noise perturbations. We show that it is feasible to establish bound analysis for evaluating noise sensitivity for SRNN by using the relation between the input current and the membrane potential change magnitude across a time window. Inspired by the theoretical analysis, we next propose a targeted penalty term in the objective function for training robust SRNN. Experimental results show that our solution outperforms the more complicated state-of-the-art methods on the commonly tested Fashion MNIST and CIFAR-IO image classification datasets.
Wendong Zheng, Yu Zhou 0029, Gang Chen 0023, Zonghua Gu 0001
ICCAD1
2023 Adaptive Optimal Electrical Resistance Tomography for Large-Area Tactile Sensing
abstract
It is critical to perceive physical contact for intelligent robots to safely interact in dynamic, unstructured environments. As physical contacts can occur at any location, a well-performing tactile sensing system should be able to deploy a large area on robotic surface. Some researchers have implemented large-area tactile sensors by using sensing arrays, but it is challenging to deploy many sensing elements. Electrical resistance tomography (ERT) has recently been introduced into tactile sensing to overcome some of the limitations with conventional tactile sensing arrays, and good results have been achieved for some robotic applications. However, a particular challenge is that spatial resolution is low. Although various attempts have been made to improve the performance of ERT-based tactile sensors, the intrinsic resolution issue remains unsolved. In this paper, we propose a novel adaptive optimal drive strategy for efficient ERT-based large-area tactile sensing for robotic applications, which can adaptively select the current injection and voltage measurement pattern for optimal tactile stimulus. In particular, regions of tactile contacts are preliminarily detected and localized by a base scanning pattern with only a few measurement data. According to this detected region, the adaptive strategy can select the optimal current injection and voltage measurement pattern to improve the sensing performance by maximizing the current density. To verify the effectiveness of the proposed strategy, the proposed method is comprehensively evaluated by simulation and experiments. The results revealed that the optimal strategy can effectively improve both spatial and temporal resolution.
Wendong Zheng, Huaping Liu 0001, Di Guo 0002, Wuqiang Yang
ICRA1
2023 A Hybrid Spiking Neurons Embedded LSTM Network for Multivariate Time Series Learning Under Concept-Drift Environment
abstract
Complicated temporal patterns can provide important information for accurate time series forecasting. Existing long short-term memory (LSTM) model with attention mechanism have achieved significant performance. However, the exponential decay of long-term memory of LSTM has not be resolved yet in these efforts, remaining a longstanding open problem in recurrent nature. This problem exhibits a bottleneck which restricts the performance of existing studies. Recently, spiking neural networks (SNNs) have shown high efficiency in capturing temporal patterns via the surrogate gradient (SG) method to resolve this issue. However, the concept-drift environment makes it impossible to pre-set the variance into the standard SG method due to time-varying data distribution. In this paper, we propose a novel adaptive and hybrid spiking (AHS) module embedded LSTM, collaborating with two attention mechanisms (called HSN-LSTM) to resolve above-mentioned problems. First, the AHS module is analyzed theoretically can remain long-term memory. Moreover, our smooth SG method avoids pre-setting of variance, which is not sensitive in the above scenarios. Besides, we use the negative log-likelihood function to adjust the attention score for alleviating the negative impact from the concept-drift. Experiment results show the HSN-LSTM outperformed the state-of-the-art models on several multivariate time series datasets.
Wendong Zheng, Putian Zhao, Gang Chen 0023, Yonghong Tian 0001
IEEE Trans. Knowl. Data Eng.1
2023 Multivariate Time Series Prediction Based on Temporal Change Information Learning Method
abstract
In the multivariate time series prediction tasks, the impact information of all nonpredictive time series on the predictive target series is difficult to be extracted at different time stages. Through the emphasis on optimal-related sequences in the target series, the deep learning model with the attention mechanism achieves a good predictive performance. However, temporal change information in the objective function and optimization algorithm is completely ignored in these models. To this end, a temporal change information learning (CIL) method is proposed in this article. First, mean absolute error (MAE) and mean squared error (MSE) losses are contained in the objective function to evaluate different amplitude errors. Meanwhile, the second-order difference technology is used in the correlation terms of the objective function to adaptively capture the impact of the abrupt and slow change information in each series on the target series. Second, the long short-term memory (LSTM) network with the transformation mechanism is used in the method so that temporal dependence information can be fully extracted (i.e., avoiding the supersaturation region). Third, to effectively obtain the optimal model parameters, the current and historical moment estimation information is adaptively memorized without the introduction of additional hyperparameters, and therefore, the acquisition ability of temporal change information in the error gradient flow is greatly enhanced by the proposed optimization algorithm. Finally, three datasets with different scales are used to verify the advantages of the CIL method in computational overhead and prediction effect.
Wendong Zheng, Jun Hu 0009
IEEE Trans. Neural Networks Learn. Syst.1
2022 An Accurate GRU-Based Power Time-Series Prediction Approach With Selective State Updating and Stochastic Optimization
abstract
Accurate power time-series prediction is an important application for building new industrialized smart cities. The gated recurrent units (GRUs) models have been successfully employed to learn temporal information for power time-series prediction, demonstrating its effectiveness. However, from a statistical perspective, these existing models are geometrically ergodic with short-term memory that causes the learned temporal information to be quickly forgotten. Meanwhile, these existing approaches completely ignore the temporal dependencies between the gradient flow in the optimization algorithm, which greatly limits the prediction accuracy. To resolve these issues, we propose a novel GRU model coupling two new mechanisms of selective state updating and adaptive mixed gradient optimization (GRU-SSU-AMG) to improve the accuracy of prediction. Specifically, a tensor discriminator is used for adaptively determining whether hidden state information needs to be updated at each time step for learning the extremely fluctuating information in the proposed selective GRU (SGRU). In addition, an adaptive mixed gradient (AdaMG) optimization method that mixes the moment estimations is proposed to further improve the capability of learning the temporal dependencies information. The effectiveness of the GRU-SSU-AMG has been extensively evaluated on five different real-world datasets. The experimental results show that the GRU-SSU-AMG achieves significant accuracy improvement compared with the state-of-the-art approaches.
Wendong Zheng, Gang Chen 0023
IEEE Trans. Cybern.1
2021 Understanding the Property of Long Term Memory for the LSTM with Attention Mechanism
abstract
Recent trends of incorporating LSTM network with different attention mechanisms in time series forecasting have led researchers to consider the attention module as an essential component. While existing studies revealed the effectiveness of attention mechanism with some visualization experiments, the underlying rationale behind their outstanding performance on learning long-term dependencies remains hitherto obscure. In this paper, we aim to elaborate on this fundamental question by conducting a thorough investigation of the memory property for LSTM network with attention mechanism. We present a theoretical analysis of LSTM integrated with attention mechanism, and demonstrate that it is capable of generating an adaptive decay rate which dynamically controls the memory decay according to the obtained attention score. In particular, our theory shows that attention mechanism brings significantly slower decays than the exponential decay rate of a standard LSTM. Experimental results on four real-world time series datasets demonstrate the superiority of the attention mechanism for maintaining long-term memory when compared to the state-of-the-art methods, and further corroborate our theoretical analysis.
Wendong Zheng, Putian Zhao, Kai Huang 0001, Gang Chen 0023
CIKM1
2021 Lifelong Visual-Tactile Cross-Modal Learning for Robotic Material Perception
abstract
The material attribute of an object's surface is critical to enable robots to perform dexterous manipulations or actively interact with their surrounding objects. Tactile sensing has shown great advantages in capturing material properties of an object's surface. However, the conventional classification method based on tactile information may not be suitable to estimate or infer material properties, particularly during interacting with unfamiliar objects in unstructured environments. Moreover, it is difficult to intuitively obtain material properties from tactile data as the tactile signals about material properties are typically dynamic time sequences. In this article, a visual-tactile cross-modal learning framework is proposed for robotic material perception. In particular, we address visual-tactile cross-modal learning in the lifelong learning setting, which is beneficial to incrementally improve the ability of robotic cross-modal material perception. To this end, we proposed a novel lifelong cross-modal learning model. Experimental results on the three publicly available data sets demonstrate the effectiveness of the proposed method.
Wendong Zheng, Huaping Liu 0001, Fuchun Sun 0001
IEEE Trans. Neural Networks Learn. Syst.1
2020 Multistage attention network for multivariate time series prediction
Jun Hu 0009, Wendong Zheng
Neurocomputing2
2020 A deep learning model to effectively capture mutation information in multivariate time series prediction
Jun Hu 0009, Wendong Zheng
Knowl. Based Syst.2
2020 Cross-Modal Material Perception for Novel Objects: A Deep Adversarial Learning Method
abstract
To more actively perform fine manipulation tasks in the real world, intelligent robots should be able to understand and communicate the physical attributes of the material during interaction with an object. Tactile and vision are two important sensing modalities in robotic perception system. In this article, we propose a cross-modal material perception framework for recognizing novel objects. Concretely, it first adopts an object-agnostic method to associate information from tactile and visual modalities. It then recognizes a novel object by using its tactile signal to retrieve perceptually similar surface material images through the learned cross-modal correlation. This problem exhibits a challenge because data from visual and tactile modalities are highly heterogeneous and weakly paired. Moreover, the framework should not only consider cross-modal pairwise relevance but also be discriminative and generalized for unseen objects. To this end, we propose a weakly paired cross-modal adversarial learning (WCMAL) model for the visual–tactile cross-modal retrieval, which combines the advantages of deep learning and adversarial learning. In particular, the model fully considers the weak pairing problem between the two modalities. Finally, we conduct verification experiments on a publicly available data set. The results demonstrate the effectiveness of the proposed method.Note to Practitioners—Since cross-modal perception can improve the active operation of automation systems, it is invaluable for industrial intelligence, particularly when only one sensing modality cannot be used or suitable in some applications. In this article, we provide a framework of cross-modal material perception for object recognition using the idea of the cross-modal retrieval. Concretely, we use relevant tactile data of an unknown object to retrieve perceptually similar surface images, which are used to evaluate its material properties. Different from that previous works using tactile information as a complement or alternative to visual information to recognize specific objects, our proposed framework is able to estimate and infer material properties of both seen and unseen objects, which can enhance manipulation systems intelligence and improve the quality of the interaction. In our future works, more modality information will be incorporated to further enhance the cross-modal material perception.
Wendong Zheng, Huaping Liu 0001, Bowen Wang 0006, Fuchun Sun 0001
IEEE Trans Autom. Sci. Eng.1
2019 Transformation-gated LSTM: efficient capture of short-term mutation dependencies for multivariate time series prediction tasks
abstract
Most multivariate time series data have very complex long-term and short-term dependencies that change over time. Currently, some recurrent neural network (RNN) variants for sequence tasks enhance the learning ability of long-term dependence on time series data. However, there lack of RNN network for capturing short-term mutation information for multivariate time series. In the present work, we proposed a transformation-gated LSTM (TG-LSTM) to enhance the ability of capturing short-term mutation information. First, the transformation gate introduced a hyperbolic tangent function to the memory cell state of the previous time step and the input gate information of the current time step without losing the memory cell state information. Then, the function value range of the partial derivative corresponding to the transformation gate during the backpropagation fully reflected the gradient change, thereby obtaining a better error gradient flow. We further extended to multi-layer TG-LSTM network and compared its stability and robustness with all baseline models. The multi-layer TG-LSTM network was superior to all baseline models in terms of prediction accuracy and performance stability on two different multivariate time series tasks.
Jun Hu 0009, Wendong Zheng
IJCNN2
2019 Cross-Modal Surface Material Retrieval Using Discriminant Adversarial Learning
abstract
The surface properties of an object play a vital role in the tasks of robotic manipulation or interaction with its surrounding environment. Tactile sensing can provide rich information about the surface properties of an object through physical contact. Hence, how to convey and interpret the tactile information to the user is a significant problem during the human–machine interaction. To this end, a visual–tactile cross-modal retrieval framework is proposed for perceptual estimation by associating tactile information to visual information of material surfaces. Namely, we can use tactile information of an unknown material surface to retrieve perceptually similar surfaces from an available surface visual sample set. For the proposed framework, we develop a discriminant adversarial learning method, which incorporates intramodal discriminant, cross-modal correlation, and intermodal consistency into a deep learning network for common feature representation learning. Experimental results on the publicly available data set show that the proposed framework and the method are effective.
Wendong Zheng, Huaping Liu 0001, Bowen Wang 0006, Fuchun Sun 0001
IEEE Trans. Ind. Informatics1