VLDB 2026 Research / reviewers in the wild / expert
Mei Tu
dblp:136/8671
· DBLP profile ↗
17ranked-venue papers
4as first author
14since 2021 · last 2024
0000-0002-4684-1759ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | A Cross Search Method for Data Augmentation in Neural Machine TranslationabstractLarge language models (LLMs) have shown excellent performance on general machine translation. However, LLMs suffer from high deployment cost and unsatisfying quality on low-resource domains. To this end, we explore to build base translation models with LLM-enhanced data augmentation. For data augmentation, we propose a cross search method to obtain qualified parallel in-domain corpus. This method encompasses two distinct approaches: antagony-cross search and similarity-cross search. Antagony-cross search helps to generate monolingual data that closely aligns with the target domain by employing token-level control. Similarity-cross search keeps the alignment between source and target sentences through a similarity score in back translation, so that the generated target language is closer to the source language semantically. With the proposed method, we generate millions of high-quality parallel in-domain corpus from low-resource monolingual data. Our proposed method achieves improvements of approximately 0.5-4 BLEU scores in these domains. Mengchao Zhang, Mei Tu |
ICASSP | 2 |
| 2023 | Pretrained Bidirectional Distillation for Machine TranslationabstractKnowledge transfer can boost neural machine translation (NMT), for example, by finetuning a pretrained masked language model (LM).However, it may suffer from the forgetting problem and the structural inconsistency between pretrained LMs and NMT models.Knowledge distillation (KD) may be a potential solution to alleviate these issues, but few studies have investigated language knowledge transfer from pretrained language models to NMT models through KD.In this paper, we propose Pretrained Bidirectional Distillation (PBD) for NMT, which aims to efficiently transfer bidirectional language knowledge from masked language pretraining to NMT models.Its advantages are reflected in efficiency and effectiveness through a globally defined and bidirectional context-aware distillation objective.Bidirectional language knowledge of the entire sequence is transferred to an NMT model concurrently during translation training.Specifically, we propose self-distilled masked language pretraining to obtain the PBD objective.We also design PBD losses to efficiently distill the language knowledge, in the form of token probabilities, to the encoder and decoder of an NMT model using the PBD objective.Extensive experiments reveal that pretrained bidirectional distillation can significantly improve machine translation performance and achieve competitive or even better results than previous pretrain-finetune or unified multilingual translation methods in supervised, unsupervised, and zero-shot scenarios.Empirically, it is concluded that pretrained bidirectional distillation is an effective and efficient method for transferring language knowledge from pretrained language models to NMT models. Yimeng Zhuang, Mei Tu |
ACL (1) | 2 |
| 2023 | Multi-Modal Beam Selection: A Transfer Methodology for Multi-FrequencyabstractThis paper investigates beam selection for multiple-frequency via deep learning. Existing learning-based beam selection methods are typically data-hungry to train the neural network. However, collecting sufficient data is a major challenge that hinders the generalization of networks. To address this challenge, this paper develops a frequency transfer method and proposes a multi-modal beam selection network, named FtransNet, which demonstrates superior generalization performance to different scenarios and carrier frequencies. Compared to the traditional approach of only transferring a model, we also transfer and augment high-quality samples to the new frequency, which significantly enlarges the environment features and path loss features in the dataset. Moreover, we embed the relative locations and the reflection features of the environment to assist the best beam selection, which further enhances the generalization of the network. Simulation results show that the proposed FtransNet outperforms the existing network with high beam selection accuracy. Dong Wang 0028, Mei Tu, Xiangfeng Gao, Zhigang Wang 0002 |
GLOBECOM | 3 |
| 2023 | A Person Identification System for the ICASSP 2023 e-Prevention ChallengeabstractThis paper describes SRCB-LUL team’s person identification system submitted to track 1 of the ICASSP 2023 Person Identification and Relapse Detection from Continuous Recordings of Biosignals (e-Prevention) challenge, which aims to identify the wearer of the smartwatch. At first, abnormal values of the multi-channel physiological signals are replaced, and the valid data are divided into multiple fine-grained fragments. Then, these fragments are fed into a 1D-CNN, and the user ID of a day is predicted through the voting of fragment results. Based on this framework, we train multiple base networks with different input lengths or different signal channels and aggregate these base networks for ensemble learning. The accuracy of our system is 96.16/95.00% on the validation/test set. Jinting Wu, Mei Tu |
ICASSP | 2 |
| 2023 | Multi-teacher Knowledge Distillation for End-to-End Text Image Machine Translation
Cong Ma 0002, Mei Tu, Yang Zhao 0007, Yu Zhou 0001, Chengqing Zong |
ICDAR (1) | 3 |
| 2023 | E2TIMT: Efficient and Effective Modal Adapter for Text Image Machine Translation
Cong Ma 0002, Mei Tu, Yang Zhao 0007, Yu Zhou 0001, Chengqing Zong |
ICDAR (6) | 3 |
| 2022 | Long-range Sequence Modeling with Predictable Sparse AttentionabstractSelf-attention mechanism has been shown to be an effective approach for capturing global context dependencies in sequence modeling, but it suffers from quadratic complexity in time and memory usage.Due to the sparsity of the attention matrix, much computation is redundant.Therefore, in this paper, we design an efficient Transformer architecture, named Fourier Sparse Attention for Transformer (FSAT), for fast long-range sequence modeling.We provide a brand-new perspective for constructing sparse attention matrix, i.e. making the sparse attention matrix predictable.Two core submodules are: (1) A fast Fourier transform based hidden state cross module, which captures and pools L 2 semantic combinations in O(L log L) time complexity.(2) A sparse attention matrix estimation module, which predicts dominant elements of an attention matrix based on the output of the previous hidden state cross module.By reparameterization and gradient truncation, FSAT successfully learned the index of dominant elements.The overall complexity about the sequence length is reduced from O(L 2 ) to O(L log L).Extensive experiments (natural language, vision, and math) show that FSAT remarkably outperforms the standard multi-head attention and its variants in various long-sequence tasks with low computational costs, and achieves new state-of-the-art results on the Long Range Arena benchmark. Yimeng Zhuang, Mei Tu |
ACL (1) | 3 |
| 2022 | A CNN Model with Discretized Mobile Features for Depression DetectionabstractDepression has been a serious mental illness for a long time, which significantly influences people’s life quality. Meanwhile, as the smartphone becomes an integral part of people’s lives, it creates the opportunity to analyze users’ feelings through their phone usage and sensor data. However, previous studies mainly adopt machine-learning methods for depression detection, ignoring the sequential patterns hidden in them. In this study, we aim to monitor the symptoms of depression through sequential mobile data collected from phones and their sensors. First, we establish a deep-learning model called Dep-caser to fully utilize the sequential information in mobile data. Next, we introduce a discretization method based on Information Value to deal with data sparsity and outliers. In total, we recruited 257 people to join the study and extracted five-day longitudinal data from their smartphones and electronic bands. We conduct two experiments to examine the effectiveness of the Dep-caser and discretization method respectively. The results demonstrate that Dep-caser outperforms most of the machine learning methods and the discretization further improves the performance of the deep-learning model to achieve an overall accuracy of 0.83. Our study shows the promising future to adopt deep-learning models with sequential phone usage and sensing data to detect depression. Yueru Yan, Mei Tu, Hongbo Wen |
BSN | 2 |
| 2022 | ASR Error Correction with Dual-Channel Self-Supervised LearningabstractTo improve the performance of Automatic Speech Recognition (ASR), it is common to deploy an error correction module at the post-processing stage to correct recognition errors. In this paper, we propose 1) an error correction model, which takes account of both contextual information and phonetic information by dual-channel; 2) a self-supervised learning method for the model. Firstly, an error region detection model is used to detect the error regions of ASR output. Then, we perform dual-channel feature extraction for the error regions, where one channel extracts their contextual information with a pre-trained language model, while the other channel builds their phonetic information. At the training stage, we construct error patterns at the phoneme level, which simplifies the data annotation procedure, thus allowing us to leverage a large scale of unlabeled data to train our model in a self-supervised learning manner. Experimental results on different test sets demonstrate the effectiveness and robustness of our model. Mei Tu, Jinyao Yan |
ICASSP | 2 |
| 2022 | Improving End-to-End Text Image Translation From the Auxiliary Text Translation TaskabstractEnd-to-end text image translation (TIT), which aims at translating the source language embedded in images to the target language, has attracted intensive attention in recent research. However, data sparsity limits the performance of end-to-end text image translation. Multi-task learning is a nontrivial way to alleviate this problem via exploring knowledge from complementary related tasks. In this paper, we propose a novel text translation enhanced text image translation, which trains the end-to-end model with text translation as an auxiliary task. By sharing model parameters and multi-task training, our model is able to take full advantage of easily-available large-scale text parallel corpus. Extensive experimental results show our proposed method outperforms existing end-to-end methods, and the joint multi-task learning with both text translation and recognition tasks achieves better results, proving translation and recognition auxiliary tasks are complementary.1 Cong Ma 0002, Mei Tu, Linghui Wu, Yang Zhao 0007, Yu Zhou 0001 |
ICPR | 3 |
| 2022 | Angular Gap: Reducing the Uncertainty of Image Difficulty through Model CalibrationabstractCurriculum learning needs example difficulty to proceed from easy to hard. However, the credibility of image difficulty is rarely investigated, which can seriously affect the effectiveness of curricula. In this work, we propose Angular Gap, a measure of difficulty based on the difference in angular distance between feature embeddings and class-weight embeddings built by hyperspherical learning. To ascertain difficulty estimation, we introduce class-wise model calibration, as a post-training technique, to the learnt hyperbolic space. This bridges the gap between probabilistic model calibration and angular distance estimation of hyperspherical learning. We show the superiority of our calibrated Angular Gap over recent difficulty metrics on CIFAR10-H and ImageNetV2. We further propose a curriculum based on Angular Gap for unsupervised domain adaptation that can translate from learning easy samples to mining hard samples. We combine this curriculum with a state-of-the-art self-training method, Cycle Self Training (CST). The proposed Curricular CST learns robust representations and outperforms recent baselines on Office31 and VisDA 2017. Bohua Peng, Mobarakol Islam, Mei Tu |
ACM Multimedia | 3 |
| 2022 | Depression Detection via Influence-based Relabeling for Resolving Training Set NoiseabstractEarly depression detection research employs machine learning models trained on crowd-sourcing data. The training data easily suffer from label noise due to weak self-perception of people and uncontrollability of collection process. The noise issue is seldom discussed in the previous depression detection work. In this work, we firstly introduce the influence-based relabeling method in the depression detection task to revise noise labels, and further move one step forward to propose a threshold ratio function to control the relabeling sample size. The relabeling sample size is usually ignored in the previous influence function, so that the relabeling is sometimes overwhelming, leading to great change on the distribution of the training data, and model performance decline. Our proposed method aims at avoiding giant change on the training data. To achieve this, we design an adjustable ratio threshold for the samples to be relabeled. The ratio is adjusted according to the trained model performance. If the model has good performance on the validation set, the relabeling ratio tends mild, otherwise, the relabeling can be aggressive. In the experiments, we recruited 205 participants and collect the usage data from smartphones and wearable bands, including participants’ response to the questionnaire Depression Anxiety Stress Scale-21. We discuss several main stream denoising methods and compare four most recent methods in the depression detection task. The proposed model achieves a best testing F1 score of 86.3%. Mei Tu, Ziang Zhao, Jinting Wu, Hwangrae Lee, Jinwoo Song |
SMC | 1 |
| 2022 | Improved Transformer With Multi-Head Dense CollaborationabstractRecently, the attention mechanism boosts the performance of many neural network models in Natural Language Processing (NLP). Among the various attention mechanisms, Multi-Head Attention (MHA) is a powerful and popular variant. MHA helps the model to attend to different feature subspaces independently which is an essential component of Transformer. Despite its success, we conjecture that the different heads of the existing MHA may not collaborate properly. To validate this assumption and further improve the performance of Transformer, we study the collaboration problem for MHA in this paper. First, we propose the Single-Layer Collaboration (SLC) mechanism to help each attention head improve its attention distribution based on the feedback of other heads. Furthermore, we extend SLC to the cross-layer Multi-Head Dense Collaboration (MHDC) mechanism. MHDC helps each MHA layer learn the attention distributions considering the knowledge from the other MHA layers. Both SLC and MHDC are implemented as lightweight modules with very few additional parameters. When equipped with these modules, our new framework, i.e., Collaborative TransFormer (CollFormer), significantly outperforms the vanilla Transformer on a range of NLP tasks, including machine translation, sentence semantic relatedness, natural language inference, sentence classification, and reading comprehension. Besides, we also carry out extensive quantitative experiments to analyze the properties of the MHDC in different settings. The experimental results validate the effectiveness and universality of MHDC as well asCollFormer. Xin Shen 0003, Mei Tu, Yimeng Zhuang, Zhiyuan Liu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | Accelerating Neural Machine Translation with Partial Word Embedding CompressionabstractLarge model size and high computational complexity prevent the neural machine translation (NMT) models from being deployed to low resource devices (e.g. mobile phones). Due to the large vocabulary, a large storage memory is required for the word embedding matrix in NMT models, in the meantime, high latency is introduced when constructing the word probability distribution. Based on reusing the word embedding matrix in the softmax layer, it is possible to handle the two problems brought by large vocabulary at the same time. In this paper, we propose Partial Vector Quantization (P-VQ) for NMT models, which can both compress the word embedding matrix and accelerate word probability prediction in the softmax layer. With P-VQ, the word embedding matrix is split into two low dimensional matrices, namely the shared part and the exclusive part. We compress the shared part by vector quantization and leave the exclusive part unchanged to maintain the uniqueness of each word. For acceleration, in the softmax layer, we replace most of the multiplication operations with the efficient looking-up operations based on our compression to reduce the computational complexity. Furthermore, we adopt curriculum learning and compact the word embedding matrix gradually to improve the compression quality. Experimental results on the Chinese-to-English translation task show that our method can reduce 74.35% of parameters of the word embedding and 74.42% of the FLOPs of the softmax layer. Meanwhile, the average BLEU score on the WMT test sets only drops 0.04. Mei Tu, Jinyao Yan |
AAAI | 2 |
| 2020 | End-to-End Speech Translation with Self-Contained Vocabulary ManipulationabstractIn machine translation, vocabulary manipulation is a way to reduce the target vocabulary based on the source sentence and the word dictionary, which is effective to lower latency during inference for text translation in industrial application. But vocabulary manipulation is hard to apply to the end-to-end speech-text translation, because neither source text nor speech-to-target mapping is available. We introduce a method that avoids this dependence. Through learning the projection between sentence-level speech encoder output and final target vocabulary, the proposed method allows self-contained vocabulary manipulation without knowing source speech transcripts or dictionaries. Experimental results show that the proposed method speed up about 20% while keep the comparable translation quality. Mei Tu |
ICASSP | 1 |
| 2015 | Exploring Diverse Features for Statistical Machine Translation Model PruningabstractIn phrase-based and hierarchical phrase-based statistical machine translation systems, translation performance depends heavily on the size and quality of the translation table. To meet the requirements of making a real-time response, some research has been performed to filter the translation table. However, most existing methods are always based on one or two constraints that act as hard rules, such as not allowing phrase-pairs with low translation probabilities. These approaches sometimes make constraints rigid because they consider only a single factor instead of composite factors. Based on the considerations above, in this paper, we propose a machine learning-based framework that integrates multiple features for translation model pruning. Experimental results show that our framework is effective by pruning 80% of the phrase-pairs and 70% of the hierarchical rules, while retaining the quality of the translation models when using the BLEU evaluation metric. Our study further shows that our method can select the most useful phrase-pairs and rules, including those that are low in frequency but still very useful. Mei Tu, Yu Zhou 0001, Chengqing Zong |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2014 | Enhancing Grammatical Cohesion: Generating Transitional Expressions for SMTabstractTransitional expressions provide glue that holds ideas together in a text and enhance the logical organization, which together help improve readability of a text. However, in most current statistical machine translation (SMT) systems, the outputs of compound-complex sentences still lack proper transitional expressions. As a result, the translations are often hard to read and understand. To address this issue, we propose two novel models to encourage generating such transitional expressions by introducing the source compoundcomplex sentence structure (CSS). Our models include a CSS-based translation model, which generates new CSS-based translation rules, and a generative transfer model, which encourages producing transitional expressions during decoding. The two models are integrated into a hierarchical phrase-based translation system to evaluate their effectiveness. The experimental results show that significant improvements are achieved on various test data meanwhile the translations are more cohesive and smooth. Mei Tu, Yu Zhou 0001, Chengqing Zong |
ACL (1) | 1 |