EDBT 2026 Demo / reviewers in the wild / expert
Weitai Zhang
dblp:160/7586
· DBLP profile ↗
11ranked-venue papers
4as first author
10since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Unlocking the Power of Patch: Patch-Based MLP for Long-Term Time Series ForecastingabstractRecent studies have attempted to refine the Transformer architecture to demonstrate its effectiveness in Long-Term Time Series Forecasting (LTSF) tasks. Despite surpassing many linear forecasting models with ever-improving performance, we remain skeptical of Transformers as a solution for LTSF. We attribute the effectiveness of these models largely to the adopted Patch mechanism, which enhances sequence locality to an extent yet fails to fully address the loss of temporal information inherent to the permutation-invariant self-attention mechanism. Further investigation suggests that simple linear layers augmented with the Patch mechanism may outperform complex Transformer-based LTSF models. Moreover, diverging from models that use channel independence, our research underscores the importance of cross-variable interactions in enhancing the performance of multivariate time series forecasting. The interaction information between variables is highly valuable but has been misapplied in past studies, leading to suboptimal cross-variable models. Based on these insights, we propose a novel and simple Patch-based MLP (PatchMLP) for LTSF tasks. Specifically, we employ simple moving averages to extract smooth components and noise-containing residuals from time series data, engaging in semantic information interchange through channel mixing and specializing in random noise with channel independence processing. The PatchMLP model consistently achieves state-of-the-art results on several real-world datasets. We hope this surprising finding will spur new research directions in the LTSF field and pave the way for more efficient and concise solutions. Peiwang Tang, Weitai Zhang |
AAAI | 2 |
| 2025 | Scalable Data Synthesis through Human-like Cognitive Imitation and Data RecombinationabstractLarge language models (LLMs) rely on massive amounts of training data, however, the quantity of empirically observed data is limited.To alleviate this issue, lots of LLMs leverage synthetic data to enhance the quantity of training data.Despite significant advancements in LLMs, the efficiency and scalability characteristics of data synthesis during pre-training phases remain insufficiently explored.In this work, we propose a novel data synthesis framework, Cognitive Combination Synthesis (CCS), designed to achieve highly efficient and scalable data synthesis.Specifically, our methodology mimics human cognitive behaviors by recombining and interconnecting heterogeneous data from diverse sources thereby enhancing advanced reasoning capabilities in LLMs.Extensive experiments demonstrate that: (1) effective data organization is essential, and our mappingbased combination learning approach significantly improves data utilization efficiency; (2) by enhancing data diversity, accuracy, and complexity, our synthetic data scales beyond 100B tokens, revealing CCS's strong scalability.Our findings highlight the impact of data organization methods on LLM learning efficiency and the significant potential of scalable synthetic data to enhance model reasoning capabilities. Zhongyi Ye, Weitai Zhang, Xinyuan Zhou, Ninghui Rao, Enhong Chen |
EMNLP | 2 |
| 2025 | Leveraging Boolean Directivity Embedding for Binaural Target Speaker ExtractionabstractDirection-based target speaker extraction (TSE) attracts a constant attention due to the convenience of direction acquisition over assistive video or enrollment audio. The direction clue heavily affects the TSE performance, which might be more seriously in the case of binaural setups due to the small-sized and irregular microphone array. In this paper, we propose Boolean Directivity Embedding (BDE) as a new direction feature in order to precisely lock onto the target speaker independent on microphone array configurations for binaural TSE (BiTSE). We design an encoder that accurately aligns the BDE with the mixed audio signals for feature fusion. Considering that the Boolean representation may contain insufficient spatial and temporal information, we enhance the BDE by incorporating the previously-proposed spatiotemporal features, showing a compatibility and stronger capacity for BiTSE. The proposed BiTSE model, which is based on the narrow-band Conformer as the backbone, can adapt to the cases of target speaker switching and moving by simply modifying the frame-wise BDE. Experimental results demonstrate the efficacy of our method in both stationary and dynamic scenarios. The proposed method is open-sourced in https://github.com/ichi131/Direction-based-BiTSE. Yichi Wang 0001, Jie Zhang 0042, Chengqian Jiang, Weitai Zhang, Zhongyi Ye, Li-Rong Dai 0001 |
ICASSP | 4 |
| 2025 | Large Language Models Are Efficient Learners as Zero-Shot Speech TranslatorsabstractSignificant progress has recently been made in combining Speech Foundation Models (SFMs) and Large Language Models (LLMs) into a unified model to tackle Speech-to-Text Translation (ST) tasks. However, fine-tuning LLMs to adapt to specific downstream tasks requires substantial resources, which is often infeasible. Therefore, this study proposes using Chain-of-Thought (CoT)-based LLMs to perform error correction on Automatic Speech Recognition results followed by translation into the target language. This approach combines SFMs and LLMs in a lightweight manner without fine-tuning or large parallel corpora. Additionally, through our proposed Translation Graph CoT (TGCoT), which involves iterative feedback and Back Translation, the model can self-check when errors are detected, effectively reducing error accumulation during the multi-step CoT reasoning process and thereby improving translation accuracy. Finally, extensive experiments across languages and models demonstrate the superiority and robustness of the proposed method. The results show that the proposed approach better unleashes LLM capabilities and adapts to downstream ST tasks with minimal resources. Chenxuan Liu, Peiwang Tang, Weitai Zhang, Sreyan Ghosh, Zhongyi Ye, Mingjia Yu |
ICASSP | 4 |
| 2025 | Adversarial Speech-Text Pre-Training for Speech TranslationabstractLarge-scale pre-training has been shown to benefit speech translation tasks. However, existing multimodal pre-training efforts rely on parallel corpora for semantic alignment, potentially limiting performance to the scale of available data and causing data imbalance. Hence, we propose an adversarial speech-text pre-training (AST) scheme for speech translation. This scheme aligns the feature distributions of speech and text modalities instead of enforcing semantic alignment based on parallel corpora. Specifically, we introduced a dual-stream mechanism that bridges speech and text modalities through speech, text, and shared encoders. In addition, we designed an adversarial bridging method that focuses on the differences in feature distributions between speech and text. Leveraging a discriminator and hidden state-level swapping strategy emphasizes semantic information in speech representations, while avoiding the limitations imposed by the scale of parallel corpora. We applied AST to both end-to-end speech translation and large model architectures. Experimental results on the IWSLT test sets demonstrate that AST improved the performance of speech translation models and is compatible with large language models. Chenxuan Liu, Weitai Zhang, Peiwang Tang, Mingjia Yu, Sreyan Ghosh, Zhongyi Ye |
ICASSP | 3 |
| 2025 | Bridging Modality Gap with Large Speech and Language Models for End-to-End Speech-to-Text TranslationabstractEnd-to-end speech-to-text translation (E2E ST) has increasingly aroused interest and attention recently, attempting to address the problem of data scarcity and modeling burden. Several attempts exploring the combination of Large Speech and Language Models into a unified model to improve E2E ST are carried out. However, the inherent differences between speech and text modalities often impede effective cross-modal and cross-lingual transfer. In this study, we introduce LaSaLM-ST, a novel model architecture built upon Pre-trained Large Speech and Language Models for improving E2E ST. Our speech encoder begins with processing the source speech sequence. An adaptor and speech decoder then project speech features into the compatible feature space for the decoder-only Large Language Model (LLM), which then aligns the representation spaces of speech and text modalities with attentive interactions. Besides, we also develop a multi-step fine-tuning method to preserve the pre-trained multilingual knowledge and keep ST fine-tuning stably. Experiments conducted on the IWSLT2023 offline ST task from English to German, Chinese and Japanese demonstrate that our methodology not only achieves state-of-the-art BLEU scores but also outperforms the highly competitive cascaded ST systems in an unrestricted setting. Weitai Zhang, Simran Naagar, Zhongyi Ye, Peiwang Tang, Xinyuan Zhou, Li-Rong Dai 0001 |
ICASSP | 1 |
| 2025 | Semi-Supervised Multilingual Alignment with Lexical Memory for Massively Parallel Text MiningabstractExisting state-of-the-art techniques that employ multilingual sentence embeddings for mining parallel texts predominantly rely on extensive supervision, which often results in sub-optimal performance in the absence of large-scale parallel training datasets. In this study, we introduce a novel method designed to extract high-quality parallel texts from monolingual corpora, particularly targeting zero-and low-resource languages. We learn language-agnostic sentence embeddings with a two-tiered training regimen: an initial phase of lexical knowledge-enhanced pretraining and a subsequent phase of supervised fine-tuning on a minimally sized parallel dataset using contrastive loss. Furthermore, we enhance the model’s performance by adopting an iterative training methodology that leverages both mined data and synthetically augmented data. We illustrate the capability of our method to create high-quality parallel text with various downstream tasks. All results suggest that the proposed method is effective and can surpass previous state-of-the-art supervised methods in zero-and low-resource scenarios. Weitai Zhang, Peiwang Tang, Simran Naagar, Zhongyi Ye |
ICASSP | 1 |
| 2024 | A Study of Multichannel Spatiotemporal Features and Knowledge Distillation on Robust Target Speaker ExtractionabstractTarget speaker extraction (TSE) based on direction of arrival (DOA) has a wide range of applications in e.g., remote conferencing, hearing aids, in-car speech interaction. Due to the inherent phase uncertainty, existing TSE methods usually suffer from speaker confusion within specific frequency bands. Imprecise DOA measurements caused by e.g., the calibration of the microphone array and ambient noises, can also deteriorate the TSE performance. In order to improve the robustness of TSE, in this work we propose several new multichannel spatiotemporal features to represent the discriminability of the target speaker. The narrow-band Conformer model is applied in combination with the proposed features to facilitate the extraction of the target speaker. In addition, we consider knowledge distillation for improving the model robustness, particularly in the presence of DOA mis-match. Experimental results on a public dataset verify the efficacy of the proposed method. Yichi Wang 0001, Jie Zhang 0042, Shihao Chen, Weitai Zhang, Zhongyi Ye, Xinyuan Zhou, Li-Rong Dai 0001 |
ICASSP | 4 |
| 2024 | Pre-Trained Acoustic-and-Textual Modeling for End-To-End Speech-To-Text TranslationabstractEnd-to-end paradigm has aroused more and more interests and attention for improving speech-to-text translation (ST) recently. Existing end-to-end models mainly attributes and attempts to address the problem of modeling burden and data scarcity, while always fail to maintain both cross-modal and cross-lingual mapping well at the same time. In this work, we investigate methods for improving endto-end ST with pre-trained acoustic-and-textual models. Our acoustic encoder and decoder begins with processing the source speech sequence as usual. A textual encoder and an adaptor module then obtain source acoustic and textual information respectively, alleviating the representation inconsistency with attentive interactions in the textual decoder. Also, we utilize pre-trained models, and develop an adaptation fine-tuning method to preserve the pre-training knowledge. Experimental results on the IWSLT2023 offline ST task from English to German, Japanese and Chinese show that our method achieves state-of-the-art BLEU scores and surpasses the strong cascaded ST counterparts in unrestricted setting. Weitai Zhang, Hanyi Zhang, Chenxuan Liu, Zhongyi Ye, Xinyuan Zhou, Li-Rong Dai 0001 |
ICASSP | 1 |
| 2023 | A Detection-based Attention Alignment Method for Document-level Neural Machine TranslationabstractPrevious works have shown that inter-sentential contextual information can lead to substantial improvements in document-level neural machine translation (DocNMT).Most existing DocNMT models focus on methods of introducing inter-sentential contextual information through attention mechanisms.Compared to intra-sentential attention, however, the long-range dependency in documentlevel attention calculation inevitably introduces meaningless contextual noise, resulting in significant performance deterioration.To address this problem, this paper proposes a detection-based attention alignment method, to help each translating word focus on relevant informative contextual words.We first introduce a context detector that automatically evaluates each source-side word's effect on the model's prediction.Based on the detection results, we align the original attention weights by integrating the cosine similarity between the aligned and original attention weights into the loss function, under a multi-task framework, which allows DocNMT to more effectively capture the document-level context.The results for three English-German (En-De) public translation datasets show that the proposed method can obtain consistent improvements over a strong G-Transformer baseline. Kang Zhong, Wu Guo, Bin Gu 0004, Weitai Zhang |
SEKE | 4 |
| 2014 | A Feature Extraction Method Based on Word Embedding for Word Similarity Computing
Weitai Zhang, Weiran Xu, Guang Chen 0003, Jun Guo 0002 |
NLPCC | 1 |