VLDB 2026 Research / reviewers in the wild / expert
Yun Tang 0002
dblp:67/764-2
· DBLP profile ↗
24ranked-venue papers
7as first author
15since 2021 · last 2023
0000-0002-3122-5881ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Hybrid Transducer and Attention based Encoder-Decoder Modeling for Speech-to-Text TasksabstractYun Tang, Anna Sun, Hirofumi Inaguma, Xinyue Chen, Ning Dong, Xutai Ma, Paden Tomasello, Juan Pino. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Yun Tang 0002, Anna Y. Sun, Hirofumi Inaguma, Xutai Ma, Paden Tomasello, Juan Pino 0001 |
ACL (1) | 1 |
| 2023 | UnitY: Two-pass Direct Speech-to-speech Translation with Discrete UnitsabstractHirofumi Inaguma, Sravya Popuri, Ilia Kulikov, Peng-Jen Chen, Changhan Wang, Yu-An Chung, Yun Tang, Ann Lee, Shinji Watanabe, Juan Pino. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Hirofumi Inaguma, Sravya Popuri, Ilia Kulikov, Peng-Jen Chen, Changhan Wang, Yu-An Chung, Yun Tang 0002, Ann Lee 0001, Shinji Watanabe 0001, Juan Pino 0001 |
ACL (1) | 7 |
| 2023 | Simple and Effective Unsupervised Speech TranslationabstractChanghan Wang, Hirofumi Inaguma, Peng-Jen Chen, Ilia Kulikov, Yun Tang, Wei-Ning Hsu, Michael Auli, Juan Pino. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Changhan Wang, Hirofumi Inaguma, Peng-Jen Chen, Ilia Kulikov, Yun Tang 0002, Wei-Ning Hsu, Michael Auli, Juan Pino 0001 |
ACL (1) | 5 |
| 2023 | Named Entity Detection and Injection for Direct Speech TranslationabstractIn a sentence, certain words are critical for its semantic. Among them, named entities (NEs) are notoriously challenging for neural models. Despite their importance, their accurate handling has been neglected in speech-to-text (S2T) translation research, and recent work has shown that S2T models perform poorly for locations and notably person names, whose spelling is challenging unless known in advance. In this work, we explore how to leverage dictionaries of NEs known to likely appear in a given context to improve S2T model outputs. Our experiments show that we can reliably detect NEs likely present in an utterance starting from S2T encoder outputs. Indeed, we demonstrate that the current detection quality is sufficient to improve NE accuracy in the translation with a 31% reduction in person name errors. Marco Gaido, Yun Tang 0002, Ilia Kulikov, Rongqing Huang, Hongyu Gong, Hirofumi Inaguma |
ICASSP | 2 |
| 2023 | Improving Speech-to-Speech Translation Through Unlabeled TextabstractDirect speech-to-speech translation (S2ST) is among the most challenging problems in the translation paradigm due to the significant scarcity of S2ST data. While effort has been made to increase the data size from unlabeled speech by cascading pretrained speech recognition (ASR), machine translation (MT) and text-to-speech (TTS) models; unlabeled text has remained relatively under-utilized to improve S2ST. We propose an effective way to utilize the massive existing unlabeled text from different languages to create a large amount of S2ST data to improve S2ST performance by applying various acoustic effects to the generated synthetic data. Empirically our method outperforms the state of the art in Spanish-English translation by up to 2 BLEU. Significant gains by the proposed method are demonstrated in extremely low-resource settings for both Spanish-English and Russian-English translations. Xuan-Phi Nguyen, Sravya Popuri, Changhan Wang, Yun Tang 0002, Ilia Kulikov, Hongyu Gong |
ICASSP | 4 |
| 2023 | Enhancing Speech-To-Speech Translation with Multiple TTS TargetsabstractIt has been known that direct speech-to-speech translation (S2ST) models usually suffer from the data scarcity issue because of the limited existing parallel materials for both source and target speech. Therefore to train a direct S2ST system, previous works usually utilize text-to-speech (TTS) systems to generate samples in the target language by augmenting the data from speech-to-text translation (S2TT). However, there is a limited investigation into how the synthesized target speech would affect the S2ST models. In this work, we analyze the effect of changing synthesized target speech for direct S2ST models. We find that simply combining the target speech from different TTS systems can potentially improve the S2ST performances. Following that, we also propose a multi-task framework that jointly optimizes the S2ST system with multiple targets from different TTS systems. Extensive experiments demonstrate that our proposed framework achieves consistent improvements (2.8 BLEU) over the baselines on the Fisher Spanish-English dataset. Jiatong Shi, Yun Tang 0002, Ann Lee 0001, Hirofumi Inaguma, Changhan Wang, Juan Pino 0001, Shinji Watanabe 0001 |
ICASSP | 2 |
| 2023 | Exploration on HuBERT with Multiple Resolution
Jiatong Shi, Yun Tang 0002, Hirofumi Inaguma, Hongyu Gong, Juan Pino 0001, Shinji Watanabe 0001 |
INTERSPEECH | 2 |
| 2022 | Direct Speech-to-Speech Translation With Discrete UnitsabstractAnn Lee, Peng-Jen Chen, Changhan Wang, Jiatao Gu, Sravya Popuri, Xutai Ma, Adam Polyak, Yossi Adi, Qing He, Yun Tang, Juan Pino, Wei-Ning Hsu. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Ann Lee 0001, Peng-Jen Chen, Changhan Wang, Jiatao Gu, Sravya Popuri, Xutai Ma, Adam Polyak, Yossi Adi, Yun Tang 0002, Juan Pino 0001, Wei-Ning Hsu |
ACL (1) | 10 |
| 2022 | Unified Speech-Text Pre-training for Speech Translation and RecognitionabstractYun Tang, Hongyu Gong, Ning Dong, Changhan Wang, Wei-Ning Hsu, Jiatao Gu, Alexei Baevski, Xian Li, Abdelrahman Mohamed, Michael Auli, Juan Pino. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Yun Tang 0002, Hongyu Gong, Changhan Wang, Wei-Ning Hsu, Jiatao Gu, Alexei Baevski, Xian Li 0003, Abdel-rahman Mohamed, Michael Auli, Juan Pino 0001 |
ACL (1) | 1 |
| 2022 | Contrastive Clustering to Mine Pseudo Parallel Data for Unsupervised Translation
Xuan-Phi Nguyen, Hongyu Gong, Yun Tang 0002, Changhan Wang, Philipp Koehn, Shafiq R. Joty |
ICLR | 3 |
| 2022 | From Start to Finish: Latency Reduction Strategies for Incremental Speech Synthesis in Simultaneous Speech-to-Speech TranslationabstractSpeech-to-speech translation (S2ST) converts input speech to speech in another language. A challenge of delivering S2ST in real time is the accumulated delay between the translation and speech synthesis modules. While recently incremental text-to-speech (iTTS) models have shown large quality improvements, they typically require additional future text inputs to reach optimal performance. In this work, we minimize the initial waiting time of iTTS by adapting the upstream speech translator to generate high-quality pseudo lookahead for the speech synthesizer. After mitigating the initial delay, we demonstrate that the duration of synthesized speech also plays a crucial role on latency. We formalize this as a latency metric and then present a simple yet effective duration-scaling approach for latency reduction. Our approaches consistently reduce latency by 0.2-0.5 second without sacrificing speech translation quality. Changhan Wang, Hongyu Gong, Xutai Ma, Yun Tang 0002, Juan Pino 0001 |
INTERSPEECH | 5 |
| 2021 | Multilingual Speech Translation from Efficient Finetuning of Pretrained ModelsabstractXian Li, Changhan Wang, Yun Tang, Chau Tran, Yuqing Tang, Juan Pino, Alexei Baevski, Alexis Conneau, Michael Auli. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Xian Li 0003, Changhan Wang, Yun Tang 0002, Chau Tran, Juan Pino 0001, Alexei Baevski, Alexis Conneau, Michael Auli |
ACL/IJCNLP (1) | 3 |
| 2021 | Improving Speech Translation by Understanding and Learning from the Auxiliary Text Translation TaskabstractYun Tang, Juan Pino, Xian Li, Changhan Wang, Dmitriy Genzel. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yun Tang 0002, Juan Pino 0001, Xian Li 0003, Changhan Wang, Dmitriy Genzel |
ACL/IJCNLP (1) | 1 |
| 2021 | A General Multi-Task Learning Framework to Leverage Text Data for Speech to Text TasksabstractAttention-based sequence-to-sequence modeling provides a powerful and elegant solution for applications that need to map one sequence to a different sequence. Its success heavily relies on the availability of large amounts of training data. This presents a challenge for speech applications where labelled speech data is very expensive to obtain, such as automatic speech recognition (ASR) and speech translation (ST). In this study, we propose a general multi-task learning framework to leverage text data for ASR and ST tasks. Two auxiliary tasks, a denoising autoencoder task and machine translation task, are proposed to be co-trained with ASR and ST tasks respectively. We demonstrate that representing text input as phoneme sequences can reduce the difference between speech and text inputs, and enhance the knowledge transfer from text corpora to the speech to text tasks. Our experiments show that the proposed method achieves a relative 10~15% word error rate reduction on the English LIBRISPEECH task compared with our baseline, and improves the speech translation quality on the MUST-C tasks by 3.6~9.2 BLEU. Yun Tang 0002, Juan Pino 0001, Changhan Wang, Xutai Ma, Dmitriy Genzel |
ICASSP | 1 |
| 2021 | Pay Better Attention to Attention: Head Selection in Multilingual and Multi-Domain Sequence ModelingabstractMulti-head attention has each of the attention heads collect salient information from different parts of an input sequence, making it a powerful mechanism for sequence modeling. Multilingual and multi-domain learning are common scenarios for sequence modeling, where the key challenge is to maximize positive transfer and mitigate negative interference across languages and domains. In this paper, we find that non-selective attention sharing is sub-optimal for achieving good generalization across all languages and domains. We further propose attention sharing strategies to facilitate parameter sharing and specialization in multilingual and multi-domain sequence modeling. Our approach automatically learns shared and specialized attention heads for different languages and domains. Evaluated in various tasks including speech recognition, text-to-text and speech-to-text translation, the proposed attention sharing strategies consistently bring gains to sequence models built upon multi-head attention. For speech-to-text translation, our approach yields an average of $+2.0$ BLEU over $13$ language directions in multilingual setting and $+2.0$ BLEU over $3$ domains in multi-domain setting. Hongyu Gong, Yun Tang 0002, Juan Pino 0001, Xian Li 0003 |
NeurIPS | 2 |
| 2020 | Zero-Shot Text-to-SQL Learning with Auxiliary TaskabstractRecent years have seen great success in the use of neural seq2seq models on the text-to-SQL task. However, little work has paid attention to how these models generalize to realistic unseen data, which naturally raises a question: does this impressive performance signify a perfect generalization model, or are there still some limitations?In this paper, we first diagnose the bottleneck of the text-to-SQL task by providing a new testbed, in which we observe that existing models present poor generalization ability on rarely-seen data. The above analysis encourages us to design a simple but effective auxiliary task, which serves as a supportive model as well as a regularization term to the generation task to increase the models' generalization. Experimentally, We evaluate our models on a large text-to-SQL dataset WikiSQL. Compared to a strong baseline coarse-to-fine model, our models improve over the baseline by more than 3% absolute in accuracy on the whole dataset. More interestingly, on a zero-shot subset test of WikiSQL, our models achieve 5% absolute accuracy gain over the baseline, clearly demonstrating its superior generalizability. Shuaichen Chang, Yun Tang 0002, Jing Huang 0019, Xiaodong He 0001, Bowen Zhou 0001 |
AAAI | 3 |
| 2020 | Orthogonal Relation Transforms with Graph Context Modeling for Knowledge Graph EmbeddingabstractDistance-based knowledge graph embeddings have shown substantial improvement on the knowledge graph link prediction task, from TransE to the latest state-of-the-art RotatE.However, complex relations such as N-to-1, 1-to-N and N-to-N still remain challenging to predict.In this work, we propose a novel distance-based approach for knowledge graph link prediction.First we extend the RotatE from 2D complex domain to high dimensional space with orthogonal transforms to model relations.The orthogonal transform embedding for relations keeps the capability for modeling symmetric/anti-symmetric, inverse and compositional relations while achieves better modeling capacity.Second, the graph context is integrated into distance scoring functions directly.Specifically, graph context is explicitly modeled via two directed context representations.Each node embedding in knowledge graph is augmented with two context representations, which are computed from the neighboring outgoing and incoming nodes/edges respectively.The proposed approach improves prediction accuracy on the difficult N-to-1, 1-to-N and N-to-N cases.Our experimental results show that it achieves state-of-the-art results on two common benchmarks FB15k-237 and WNRR-18, especially on FB15k-237 which has many high in-degree nodes.Code available at https://github. com/JD-AI-Research-Silicon-Valley/ KGEmbedding-OTE. Yun Tang 0002, Jing Huang 0019, Guangtao Wang, Xiaodong He 0001, Bowen Zhou 0001 |
ACL | 1 |
| 2020 | Self-Training for End-to-End Speech TranslationabstractOne of the main challenges for end-to-end speech translation is data scarcity.We leverage pseudo-labels generated from unlabeled audio by a cascade and an end-to-end speech translation model.This provides 8.3 and 5.7 BLEU gains over a strong semi-supervised baseline on the MuST-C English-French and English-German datasets, reaching state-of-the art performance.The effect of the quality of the pseudo-labels is investigated.Our approach is shown to be more effective than simply pre-training the encoder on the speech recognition task.Finally, we demonstrate the effectiveness of self-training by directly generating pseudo-labels with an end-to-end model instead of a cascade model. Juan Pino 0001, Qiantong Xu, Xutai Ma, Mohammad Javad Dousti, Yun Tang 0002 |
INTERSPEECH | 5 |
| 2019 | End-to-End Structure-Aware Convolutional Networks for Knowledge Base CompletionabstractKnowledge graph embedding has been an active research topic for knowledge base completion, with progressive improvement from the initial TransE, TransH, DistMult et al to the current state-of-the-art ConvE. ConvE uses 2D convolution over embeddings and multiple layers of nonlinear features to model knowledge graphs. The model can be efficiently trained and scalable to large knowledge graphs. However, there is no structure enforcement in the embedding space of ConvE. The recent graph convolutional network (GCN) provides another way of learning graph node embedding by successfully utilizing graph connectivity structure. In this work, we propose a novel end-to-end StructureAware Convolutional Network (SACN) that takes the benefit of GCN and ConvE together. SACN consists of an encoder of a weighted graph convolutional network (WGCN), and a decoder of a convolutional network called Conv-TransE. WGCN utilizes knowledge graph node structure, node attributes and edge relation types. It has learnable weights that adapt the amount of information from neighbors used in local aggregation, leading to more accurate embeddings of graph nodes. Node attributes in the graph are represented as additional nodes in the WGCN. The decoder Conv-TransE enables the state-of-the-art ConvE to be translational between entities and relations while keeps the same link prediction performance as ConvE. We demonstrate the effectiveness of the proposed SACN on standard FB15k-237 and WN18RR datasets, and it gives about 10% relative improvement over the state-of-theart ConvE in terms of HITS@1, HITS@3 and HITS@10. Yun Tang 0002, Jing Huang 0019, Jinbo Bi, Xiaodong He 0001, Bowen Zhou 0001 |
AAAI | 2 |
| 2019 | Multi-hop Reading Comprehension across Multiple Documents by Reasoning over Heterogeneous GraphsabstractMulti-hop reading comprehension (RC) across documents poses new challenge over single-document RC because it requires reasoning over multiple documents to reach the final answer. In this paper, we propose a new model to tackle the multi-hop RC problem. We introduce a heterogeneous graph with different types of nodes and edges, which is named as Heterogeneous Document-Entity (HDE) graph. The advantage of HDE graph is that it contains different granularity levels of information including candidates, documents and entities in specific document contexts. Our proposed model can do reasoning over the HDE graph with nodes representation initialized with co-attention and self-attention based context encoders. We employ Graph Neural Networks (GNN) based message passing algorithms to accumulate evidences on the proposed HDE graph. Evaluated on the blind test set of the Qangaroo WikiHop data set, our HDE graph based single model delivers competitive result, and the ensemble model achieves the state-of-the-art performance. Guangtao Wang, Jing Huang 0019, Yun Tang 0002, Xiaodong He 0001, Bowen Zhou 0001 |
ACL (1) | 4 |
| 2019 | Relation Module for Non-Answerable Predictions on Reading ComprehensionabstractMachine reading comprehension (MRC) has attracted significant amounts of research attention recently, due to an increase of challenging reading comprehension datasets.In this paper, we aim to improve a MRC model's ability to determine whether a question has an answer in a given context (e.g. the recently proposed SQuAD 2.0 task).Our solution is a relation module that is adaptable to any MRC model.The relation module consists of both semantic extraction and relational information.We first extract high level semantics as objects from both question and context with multihead self-attentive pooling.These semantic objects are then passed to a relation network, which generates relationship scores for each object pair in a sentence.These scores are used to determine whether a question is nonanswerable.We test the relation module on the SQuAD 2.0 dataset using both the BiDAF and BERT models as baseline readers.We obtain 1.8% gain of F1 accuracy on top of the BiDAF reader, and 1.0% on top of the BERT base model.These results show the effectiveness of our relation module on MRC. Kevin Huang 0002, Yun Tang 0002, Jing Huang 0019, Xiaodong He 0001, Bowen Zhou 0001 |
CoNLL | 2 |
| 2019 | Deep Speaker Embedding Learning with Multi-level Pooling for Text-independent Speaker VerificationabstractThis paper aims to improve the widely used deep speaker embedding x-vector model. We propose the following improvements: (1) a hybrid neural network structure using both time delay neural network (TDNN) and long short-term memory neural networks (LSTM) to generate complementary speaker information at different levels; (2) a multi-level pooling strategy to collect speaker information from both TDNN and LSTM layers; (3) a regularization scheme on the speaker embedding extraction layer to make the extracted embeddings suitable for the following fusion step. The synergy of these improvements are shown on the NIST SRE 2016 eval test (with a 19% EER reduction) and SRE 2018 dev test (with a 9% EER reduction), as well as more than 10% DCF scores reduction on these two test sets over the x-vector baseline. Yun Tang 0002, Guo-Hong Ding, Jing Huang 0019, Xiaodong He 0001, Bowen Zhou 0001 |
ICASSP | 1 |
| 2019 | Multi-Stride Self-Attention for Speech Recognition
Kyu Jeong Han, Jing Huang 0019, Yun Tang 0002, Xiaodong He 0001, Bowen Zhou 0001 |
INTERSPEECH | 3 |
| 2006 | One-Pass Coarse-to-Fine Segmental Speech Decoding AlgorithmabstractIn this paper, a novel one-pass coarse-to-fine decoding algorithm is proposed to accelerate the speed of segment model (SM). The algorithm is originated from the segmentation similarity observation described in the paper and is specific for the SM based speech recognition. At each step, a coarse search is first implemented to get coarse segmentations and then a fine search is performed based on the derived segmentation information. This fast algorithm is successfully integrated into an SM based Mandarin LVCSR system and saves more than 50% decoding time without obvious influence on the recognition accuracy Yun Tang 0002, Hua Zhang 0009, Bo Xu 0002, Guo-Hong Ding |
ICASSP (1) | 1 |