Xi Ai

dblp:294/0314 · DBLP profile ↗
← Back
8ranked-venue papers
6as first author
8since 2021 · last 2024
0000-0002-4241-3837ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 6 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Machine translation · 34% Deep learning architectures and training · 27% Information extraction and text analysis · 14%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Machine translation › unsupervised machine translation
unsupervised neural machine translation
1.222023
On-the-fly Cross-lingual Masking for Multilingual Pre-training · ACL (1) 2023
Empirical Regularization for Synthetic Sentence Pairs in Unsupervised Neural Machine Translation · AAAI 2021
Natural language and speech › Speech recognition and synthesis › visual speech recognition
lip reading
0.712023
Cross-Modal Language Modeling in Multi-Motion-Informed Context for Lip Reading · IEEE ACM Trans. Audio Speech Lang. Process. 2023
Natural language and speech › Information extraction and text analysis › multilingual NLP › multilingual language modeling
multilingual pretraining
0.712023
On-the-fly Cross-lingual Masking for Multilingual Pre-training · ACL (1) 2023
Machine learning › Generative modeling
iterative refinement
0.612022
Leveraging Relaxed Equilibrium by Lazy Transition for Sequence Modeling · ACL (1) 2022
Machine learning › Deep learning architectures and training
sequence modeling
0.612022
Leveraging Relaxed Equilibrium by Lazy Transition for Sequence Modeling · ACL (1) 2022
Machine learning › Deep learning architectures and training
transformer
0.612022
Leveraging Relaxed Equilibrium by Lazy Transition for Sequence Modeling · ACL (1) 2022
Machine learning › Deep learning architectures and training › transformer › recurrent transformer
universal transformer
0.612022
Leveraging Relaxed Equilibrium by Lazy Transition for Sequence Modeling · ACL (1) 2022
Natural language and speech › Machine translation › monolingual data augmentation
back-translation
0.512021
Empirical Regularization for Synthetic Sentence Pairs in Unsupervised Neural Machine Translation · AAAI 2021
Natural language and speech › Machine translation
synthetic parallel data
0.512021
Empirical Regularization for Synthetic Sentence Pairs in Unsupervised Neural Machine Translation · AAAI 2021
Natural language and speech › Speech recognition and synthesis
automatic speech recognition
0.212023
Cross-Modal Language Modeling in Multi-Motion-Informed Context for Lip Reading · IEEE ACM Trans. Audio Speech Lang. Process. 2023
Natural language and speech › Information extraction and text analysis › multilingual NLP
cross-lingual classification
0.212023
On-the-fly Cross-lingual Masking for Multilingual Pre-training · ACL (1) 2023

Methods — techniques the papers use, named apart from their topics

source-target attention · 0.7piece-wise pre-training · 0.7multi-task learning · 0.7masked language modeling · 0.7cross-lingual prototype · 0.7lazy transition · 0.6equilibrium · 0.6regularization loss · 0.5joint training · 0.5back-translation · 0.5
YearPublicationVenuePosition
2024 Perceiving Multi-Layer Representations for No-reference Image Quality Assessment
abstract
In this paper, we propose an end-to-end no-reference image quality assessment (NR-IQA) method that perceives multilayer representations from low-level to high-level stages. First, multi-layer representations (MR) of the distorted images are extracted from different layers of multiple feature extraction networks to obtain fine-grained information. Second, a gated recurrent unit (GRU)-based fusion encoder (GFE) is presented to model the interrelationships between multi-layer representations, thereby generating the global feature. Finally, we construct a perception-oriented quality regression network (PQRN) to generate the quality scores. Experimental results on commonly used benchmark datasets verify the effectiveness of our proposed method over existing state-of-the-art approaches by a large margin.
Qunyue Huang, Bin Fang 0001, Xi Ai, Tianyu Nie
ICASSP3
2023 On-the-fly Cross-lingual Masking for Multilingual Pre-training
abstract
In multilingual pre-training with the objective of MLM (masked language modeling) on multiple monolingual corpora, multilingual models only learn cross-linguality implicitly from isomorphic spaces formed by overlapping different language spaces due to the lack of explicit cross-lingual forward pass.In this work, we present CLPM (Cross-lingual Prototype Masking), a dynamic and token-wise masking scheme, for multilingual pre-training, using a special token [C] x to replace a random token x in the input sentence.[C] x is a cross-lingual prototype for x and then forms an explicit crosslingual forward pass.We instantiate CLPM for the multilingual pre-training phase of UNMT (unsupervised neural machine translation), and experiments show that CLPM can consistently improve the performance of UNMT models on {De, Ro, N e} ↔ En.Beyond UNMT or bilingual tasks, we show that CLPM can consistently improve the performance of multilingual models on cross-lingual classification.
Xi Ai, Bin Fang 0001
ACL (1)1
2023 A Global-Local Contrastive Learning Framework for Video Captioning
abstract
In this paper, a global-local contrastive learning framework is proposed to leverage global contextual information from different modalities and then effectively fuse them with the supervision of contrastive learning. First, a global-local encoder is proposed to sufficiently explore the salient contextual information from different modalities, which generates the global contextual information. Second, contrastive learning is used to minimize the semantic distance between the paired modalities, which can improve the content matching between videos and the predicted captions. Finally, an attention-based multimodal encoder is presented to effectively fuse different modalities, thereby generating the multimodal representations that include global contextual information from different modalities. Extensive experimental results on benchmark datasets indicate that our proposed method is superior to the state-of-the-art approaches.
Qunyue Huang, Bin Fang 0001, Xi Ai
ICIP3
2023 Cross-Modal Language Modeling in Multi-Motion-Informed Context for Lip Reading
abstract
We observe that for lip reading, the language is locally transformed, instead of globally transformed, i.e., speaking and writing follow the same basic grammar rules. In this work, we present a cross-modal language model to tackle the lip-reading challenge on silent videos. Compared to previous works, we consider multi-motion-informed contexts composed of multiple lip-motion representations from different subspaces to guide decoding via the source-target attention mechanism. We present a piece-wise pre-training strategy inspired by multi-task learning to pre-train a visual module to generate multi-motioninformed contexts for cross-modality and pre-train a decoder to generate texts for language modeling. Our final large-scale model outperforms baseline models on four datasets: LRS2, LRS3, LRW, and GRID. We will open our source code on GitHub.
Xi Ai, Bin Fang 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2022 Leveraging Relaxed Equilibrium by Lazy Transition for Sequence Modeling
abstract
In sequence modeling, certain tokens are usually less ambiguous than others, and representations of these tokens require fewer refinements for disambiguation.However, given the nature of attention-based models like Transformer and UT (universal transformer), all tokens are equally processed towards depth.Inspired by the equilibrium phenomenon, we present a lazy transition, a mechanism to adjust the significance of iterative refinements for each token representation.Our lazy transition is deployed on top of UT to build LT (lazy transformer), where all tokens are processed unequally towards depth.Eventually, LT is encouraged to oscillate around a relaxed equilibrium.Our experiments show that LT outperforms baseline models on several tasks of machine translation, pre-training, Learning to Execute, and LAMBADA.
Xi Ai, Bin Fang 0001
ACL (1)1
2022 Vocabulary-informed Language Encoding
abstract
A Multilingual model relies on language encodings to identify input languages because the multilingual model has to distinguish between the input and output languages or among all the languages for cross-lingual tasks. Furthermore, we find that language encodings potentially refine multiple morphologies of different languages to form a better isomorphic space for multilinguality. To leverage this observation, we present a method to compute a vocabulary-informed language encoding as the language representation, for a required language, considering a local vocabulary covering an acceptable amount of the most frequent word embeddings in this language. In our experiments, our method can consistently improve the performance of multilingual models on unsupervised neural machine translation and cross-lingual embedding.
Xi Ai, Bin Fang 0001
COLING1
2021 Empirical Regularization for Synthetic Sentence Pairs in Unsupervised Neural Machine Translation
abstract
UNMT tackles translation on monolingual corpora in two required languages. Since there is no explicitly cross-lingual signal, pre-training and synthetic sentence pairs are significant to the success of UNMT. In this work, we empirically study the core training procedure of UNMT to analyze the synthetic sentence pairs obtained from back-translation. We introduce new losses to UNMT to regularize the synthetic sentence pairs by jointly training the UNMT objective and the regularization objective. Our comprehensive experiments support that our method can generally improve the performance of currently successful models on three similar pairs {French, German, Romanian} English and one dissimilar pair Russian English with acceptably additional cost.
Xi Ai, Bin Fang 0001
AAAI1
2021 Almost Free Semantic Draft for Neural Machine Translation
abstract
Translation quality can be improved by global information from the required target sentence because the decoder can understand both past and future information.However, the model needs additional cost to produce and consider such global information.In this work, to inject global information but also save cost, we present an efficient method to sample and consider a semantic draft as global information from semantic space for decoding with almost free of cost.Unlike other successful adaptations, we do not have to perform an EM-like process that repeatedly samples a possible semantic from the semantic space.Empirical experiments show that the presented method can achieve competitive performance in common language pairs with a clear advantage in inference efficiency.We will open all our source code on GitHub.
Xi Ai, Bin Fang 0001
NAACL-HLT1