Peiwang Tang

dblp:329/3841 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
9since 2021 · last 2025
0000-0003-3681-0761ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Unlocking the Power of Patch: Patch-Based MLP for Long-Term Time Series Forecasting
abstract
Recent studies have attempted to refine the Transformer architecture to demonstrate its effectiveness in Long-Term Time Series Forecasting (LTSF) tasks. Despite surpassing many linear forecasting models with ever-improving performance, we remain skeptical of Transformers as a solution for LTSF. We attribute the effectiveness of these models largely to the adopted Patch mechanism, which enhances sequence locality to an extent yet fails to fully address the loss of temporal information inherent to the permutation-invariant self-attention mechanism. Further investigation suggests that simple linear layers augmented with the Patch mechanism may outperform complex Transformer-based LTSF models. Moreover, diverging from models that use channel independence, our research underscores the importance of cross-variable interactions in enhancing the performance of multivariate time series forecasting. The interaction information between variables is highly valuable but has been misapplied in past studies, leading to suboptimal cross-variable models. Based on these insights, we propose a novel and simple Patch-based MLP (PatchMLP) for LTSF tasks. Specifically, we employ simple moving averages to extract smooth components and noise-containing residuals from time series data, engaging in semantic information interchange through channel mixing and specializing in random noise with channel independence processing. The PatchMLP model consistently achieves state-of-the-art results on several real-world datasets. We hope this surprising finding will spur new research directions in the LTSF field and pave the way for more efficient and concise solutions.
Peiwang Tang, Weitai Zhang
AAAI1
2025 Large Language Models Are Efficient Learners as Zero-Shot Speech Translators
abstract
Significant progress has recently been made in combining Speech Foundation Models (SFMs) and Large Language Models (LLMs) into a unified model to tackle Speech-to-Text Translation (ST) tasks. However, fine-tuning LLMs to adapt to specific downstream tasks requires substantial resources, which is often infeasible. Therefore, this study proposes using Chain-of-Thought (CoT)-based LLMs to perform error correction on Automatic Speech Recognition results followed by translation into the target language. This approach combines SFMs and LLMs in a lightweight manner without fine-tuning or large parallel corpora. Additionally, through our proposed Translation Graph CoT (TGCoT), which involves iterative feedback and Back Translation, the model can self-check when errors are detected, effectively reducing error accumulation during the multi-step CoT reasoning process and thereby improving translation accuracy. Finally, extensive experiments across languages and models demonstrate the superiority and robustness of the proposed method. The results show that the proposed approach better unleashes LLM capabilities and adapts to downstream ST tasks with minimal resources.
Chenxuan Liu, Peiwang Tang, Weitai Zhang, Sreyan Ghosh, Zhongyi Ye, Mingjia Yu
ICASSP3
2025 Adversarial Speech-Text Pre-Training for Speech Translation
abstract
Large-scale pre-training has been shown to benefit speech translation tasks. However, existing multimodal pre-training efforts rely on parallel corpora for semantic alignment, potentially limiting performance to the scale of available data and causing data imbalance. Hence, we propose an adversarial speech-text pre-training (AST) scheme for speech translation. This scheme aligns the feature distributions of speech and text modalities instead of enforcing semantic alignment based on parallel corpora. Specifically, we introduced a dual-stream mechanism that bridges speech and text modalities through speech, text, and shared encoders. In addition, we designed an adversarial bridging method that focuses on the differences in feature distributions between speech and text. Leveraging a discriminator and hidden state-level swapping strategy emphasizes semantic information in speech representations, while avoiding the limitations imposed by the scale of parallel corpora. We applied AST to both end-to-end speech translation and large model architectures. Experimental results on the IWSLT test sets demonstrate that AST improved the performance of speech translation models and is compatible with large language models.
Chenxuan Liu, Weitai Zhang, Peiwang Tang, Mingjia Yu, Sreyan Ghosh, Zhongyi Ye
ICASSP5
2025 Bridging Modality Gap with Large Speech and Language Models for End-to-End Speech-to-Text Translation
abstract
End-to-end speech-to-text translation (E2E ST) has increasingly aroused interest and attention recently, attempting to address the problem of data scarcity and modeling burden. Several attempts exploring the combination of Large Speech and Language Models into a unified model to improve E2E ST are carried out. However, the inherent differences between speech and text modalities often impede effective cross-modal and cross-lingual transfer. In this study, we introduce LaSaLM-ST, a novel model architecture built upon Pre-trained Large Speech and Language Models for improving E2E ST. Our speech encoder begins with processing the source speech sequence. An adaptor and speech decoder then project speech features into the compatible feature space for the decoder-only Large Language Model (LLM), which then aligns the representation spaces of speech and text modalities with attentive interactions. Besides, we also develop a multi-step fine-tuning method to preserve the pre-trained multilingual knowledge and keep ST fine-tuning stably. Experiments conducted on the IWSLT2023 offline ST task from English to German, Chinese and Japanese demonstrate that our methodology not only achieves state-of-the-art BLEU scores but also outperforms the highly competitive cascaded ST systems in an unrestricted setting.
Weitai Zhang, Simran Naagar, Zhongyi Ye, Peiwang Tang, Xinyuan Zhou, Li-Rong Dai 0001
ICASSP4
2025 Semi-Supervised Multilingual Alignment with Lexical Memory for Massively Parallel Text Mining
abstract
Existing state-of-the-art techniques that employ multilingual sentence embeddings for mining parallel texts predominantly rely on extensive supervision, which often results in sub-optimal performance in the absence of large-scale parallel training datasets. In this study, we introduce a novel method designed to extract high-quality parallel texts from monolingual corpora, particularly targeting zero-and low-resource languages. We learn language-agnostic sentence embeddings with a two-tiered training regimen: an initial phase of lexical knowledge-enhanced pretraining and a subsequent phase of supervised fine-tuning on a minimally sized parallel dataset using contrastive loss. Furthermore, we enhance the model’s performance by adopting an iterative training methodology that leverages both mined data and synthetically augmented data. We illustrate the capability of our method to create high-quality parallel text with various downstream tasks. All results suggest that the proposed method is effective and can surpass previous state-of-the-art supervised methods in zero-and low-resource scenarios.
Weitai Zhang, Peiwang Tang, Simran Naagar, Zhongyi Ye
ICASSP2
2023 Denoise Enhanced Neural Network with Efficient Data Generation for Automatic Sleep Stage Classification of Class Imbalance
abstract
Signal-channel electroencephalogram (EEG) based automatic sleep stage classification with machine learning is widely used in the study of sleep quality and analysis of sleep disorders. While due to the inevitable class imbalance problem, relatively poor accuracy in the detection of the first stage of the Non-Rapid-Eye-Movement (N1) is found. In this work, we propose to use a generative adversarial network (GAN) based on transformer encoder to reduce the class imbalance problem. In order to ensure the quality of the synthetic signal, the percentage form of the Kullback-Leibler divergence (KLP) index is designed to measure the similarity of synthetic signals generated by GAN and real ones. Meanwhile, we design a Residual Shrinkage Sequence Network for Sleep Staging (RsSleepNet) as the baseline to compare other resolutions of class imbalance with ours. The performance of the new method with the combination of GAN and RsSleepNet is effectively verified on two public datasets from PhysioNet, in which the accuracy of the N1 stage can be improved by more than 10% as compared to the current state-of-the-art approaches, largely alleviating the class imbalance problem in automatic sleep stage classification.
Peiwang Tang, Xianchao Zhang 0002
IJCNN2
2023 A Recurrent Neural Network based Generative Adversarial Network for Long Multivariate Time Series Forecasting
abstract
Some multimedia data from real life can be collected as multivariate time series data, such as community-contributed social data or sensor data. Many methods have been proposed for multivariate time series forecasting. In light of its importance in wide applications including traffic or electric power forecasting, appearance of the Transformer model has rapidly revolutionized various architectural design efforts. In Transformer, self-attention is used to achieve state-of-the-art prediction, and further studied for time series modeling in the frequency recently. These related works prove that self-attention mechanisms can reach a satisfied performance whether in time or frequency domain, but we used recurrent neural network (RNN) to verify that these are not critical and necessary. The correlation structure of RNN has time series specific inductive bias, but there are still some shortcomings in long multivariate time series forecasting. To break the forecasting bottleneck of traditional RNN architectures, we introduced RNNGAN, a novel and competitive RNN-based architecture combining the generation capability of Generative Adversarial Network (GAN) with the forecasting power of RNN. Differentiated from the Transformer, RNNGAN uses long short-term memory (LSTM) instead of the self-attention layers to model long-range dependencies. The experiment shows that, compared with the state-of-the-art models, RNNGAN can obtain competitive scores in many benchmark tests when training on multivariate time series datasets in many different fields.
Peiwang Tang, Xianchao Zhang 0002
ICMR1
2022 MTSMAE: Masked Autoencoders for Multivariate Time-Series Forecasting
abstract
Large-scale self-supervised pre-training Transformer architecture have significantly boosted the performance for various tasks in natural language processing (NLP) and computer vision (CV). However, there is a lack of researches on processing multivariate time-series by pre-trained Transformer, and especially, current study on masking time-series for self-supervised learning is still a gap. Different from language and image processing, the information density of time-series increases the difficulty of research. The challenge goes further with the invalidity of the previous patch embedding and mask methods. In this paper, according to the data characteristics of multivariate time-series, a patch embedding method is proposed, and we present an self-supervised pre-training approach based on Masked Autoencoders (MAE), called MTSMAE, which can improve the performance significantly over supervised learning without pre-training. Evaluating our method on several common multivariate time-series datasets from different fields and with different characteristics, experiment results demonstrate that the performance of our method is significantly better than the best method currently available.
Peiwang Tang, Xianchao Zhang 0002
ICTAI1
2022 Features Fusion Framework for Multimodal Irregular Time-series Events
Peiwang Tang, Xianchao Zhang 0002
PRICAI (1)1