VLDB 2026 Research / reviewers in the wild / expert
Xing Niu 0001
dblp:87/9555
· DBLP profile ↗
17ranked-venue papers
7as first author
6since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 5 first-author · 4 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Machine translation · 44% Efficient and distributed learning · 27% Language models and text generation · 10% |
Topics — the 12 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Machine translation
neural machine translation |
1.1 | 2 | 2023 | Pseudo-label Training and Model Inertia in Neural Machine Translation · ICLR 2023 Evaluating Robustness to Input Perturbations for Neural Machine Translation · ACL 2020 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.9 | 1 | 2025 | Effective post-training embedding compression via temperature control in contrastive training · ICLR 2025 |
Machine learning › Efficient and distributed learning › model compression
embedding compression |
0.9 | 1 | 2025 | Effective post-training embedding compression via temperature control in contrastive training · ICLR 2025 |
Machine learning › Efficient and distributed learning
model compression |
0.9 | 1 | 2025 | Effective post-training embedding compression via temperature control in contrastive training · ICLR 2025 |
Natural language and speech › Machine translation › controllable machine translation
formality control |
0.7 | 2 | 2020 | Controlling Neural Machine Translation Formality with Synthetic Supervision · AAAI 2020 A Study of Style in Machine Translation: Controlling the Formality of Machine Translation Output · EMNLP 2017 |
Natural language and speech › Machine translation
speech translation |
0.7 | 1 | 2023 | End-to-End Single-Channel Speaker-Turn Aware Conversational Speech Translation · EMNLP 2023 |
Natural language and speech › Machine translation
machine translation evaluation |
0.6 | 1 | 2022 | MT-GenEval: A Counterfactual and Contextual Dataset for Evaluating Gender Accuracy in Machine Translation · EMNLP 2022 |
Natural language and speech › Language models and text generation
controllable text generation |
0.4 | 1 | 2020 | Controlling Neural Machine Translation Formality with Synthetic Supervision · AAAI 2020 |
Natural language and speech › Language models and text generation › tokenization › subword tokenization
subword regularization |
0.4 | 1 | 2020 | Evaluating Robustness to Input Perturbations for Neural Machine Translation · ACL 2020 |
Natural language and speech › Machine translation
controllable machine translation |
0.3 | 1 | 2017 | A Study of Style in Machine Translation: Controlling the Formality of Machine Translation Output · EMNLP 2017 |
Machine learning › Trustworthy machine learning
fairness |
0.2 | 1 | 2022 | MT-GenEval: A Counterfactual and Contextual Dataset for Evaluating Gender Accuracy in Machine Translation · EMNLP 2022 |
Natural language and speech › Machine translation
gender bias in machine translation |
0.2 | 1 | 2022 | MT-GenEval: A Counterfactual and Contextual Dataset for Evaluating Gender Accuracy in Machine Translation · EMNLP 2022 |
Methods — techniques the papers use, named apart from their topics
contrastive learning · 0.9pseudo-labeling · 0.7end-to-end model · 0.7counterfactual evaluation · 0.6contextual evaluation · 0.6synthetic supervision · 0.4sequence-to-sequence model · 0.4robustness metrics · 0.4multi-task learning · 0.4BLEU · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech TranslationabstractAudio-Visual Speech-to-Speech Translation (AVS2S) typically prioritizes improving translation quality and naturalness. However, an equally critical aspect in audio-visual content is lip-synchrony—ensuring that the movements of the lips match the spoken content—essential for maintaining realism in dubbed videos. Despite its importance, the inclusion of lip-synchrony constraints in AVS2S models has been largely overlooked. This study addresses this gap by integrating a lip-synchrony loss into the training process of AVS2S models. Our proposed method significantly enhances lip-synchrony in direct audio-visual speechto-speech translation, achieving an average LSE-D score of 10.67, representing a 9.2% reduction in LSE-D over a strong baseline across four language pairs. Additionally, it maintains the naturalness and high quality of the translated speech when overlaid onto the original video, without any degradation in translation quality. Lucas Goncalves, Prashant Mathur, Xing Niu 0001, Chandrashekhar Lavania, Brady Houston, Srikanth Vishnubhotla, Lijia Sun, Anthony Ferritto |
ICASSP | 3 |
| 2025 | Zero-resource Speech Translation and Recognition with LLMsabstractDespite recent advancements in speech processing, zero-resource speech translation (ST) and automatic speech recognition (ASR) remain challenging problems. In this work, we propose to leverage a multilingual Large Language Model (LLM) to perform ST and ASR in languages for which the model has never seen paired audio-text data. We achieve this by using a pre-trained multilingual speech encoder, a multilingual LLM, and a lightweight adaptation module that maps the audio representations to the token embedding space of the LLM. We perform several experiments both in ST and ASR to understand how to best train the model and what data has the most impact on performance in previously unseen languages. In ST, our best model is capable to achieve BLEU scores over 23 in CoVoST2 for two previously unseen languages, while in ASR, we achieve WERs of up to 28.2%. We finally show that the performance of our system is bounded by the ability of the LLM to output text in the desired language. Karel Mundnich, Xing Niu 0001, Prashant Mathur, Srikanth Ronanki, Brady Houston, Veera Raghavendra Elluru, Nilaksh Das, Zejiang Hou, Goeric Huybrechts, Anshu Bhatia, Daniel Garcia-Romero, Kyu J. Han, Katrin Kirchhoff |
ICASSP | 2 |
| 2025 | Effective post-training embedding compression via temperature control in contrastive trainingabstractFixed-size learned representations (dense representations, or embeddings) are widely used in many machine learning applications across language, vision or speech modalities. This paper investigates the role of the temperature parameter in contrastive training for text embeddings. We shed light on the impact this parameter has on the intrinsic dimensionality of the embedding spaces obtained, and show that lower intrinsic dimensionality is further correlated with effective compression of embeddings. We still observe a trade-off between absolute performance and effective compression and we propose temperature aggregation methods which reduce embedding size by an order of magnitude with minimal impact on quality. Georgiana Dinu, Corey D. Barrett, Miguel Romero Calvo, Anna Currey, Xing Niu 0001 |
ICLR | 6 |
| 2023 | End-to-End Single-Channel Speaker-Turn Aware Conversational Speech TranslationabstractJuan Pablo Zuluaga-Gomez, Zhaocheng Huang, Xing Niu, Rohit Paturi, Sundararajan Srinivasan, Prashant Mathur, Brian Thompson, Marcello Federico. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Juan Zuluaga-Gomez, Zhaocheng Huang, Xing Niu 0001, Rohit Paturi, Sundararajan Srinivasan, Prashant Mathur, Brian Thompson 0001, Marcello Federico |
EMNLP | 3 |
| 2023 | Pseudo-label Training and Model Inertia in Neural Machine Translation
Benjamin Hsu, Anna Currey, Xing Niu 0001, Maria Nadejde, Georgiana Dinu |
ICLR | 3 |
| 2022 | MT-GenEval: A Counterfactual and Contextual Dataset for Evaluating Gender Accuracy in Machine TranslationabstractAnna Currey, Maria Nadejde, Raghavendra Reddy Pappagari, Mia Mayer, Stanislas Lauly, Xing Niu, Benjamin Hsu, Georgiana Dinu. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Anna Currey, Maria Nadejde, Raghavendra Reddy Pappagari, Mia Mayer, Stanislas Lauly, Xing Niu 0001, Benjamin Hsu, Georgiana Dinu |
EMNLP | 6 |
| 2020 | Controlling Neural Machine Translation Formality with Synthetic SupervisionabstractThis work aims to produce translations that convey source language content at a formality level that is appropriate for a particular audience. Framing this problem as a neural sequence-to-sequence task ideally requires training triplets consisting of a bilingual sentence pair labeled with target language formality. However, in practice, available training examples are limited to English sentence pairs of different styles, and bilingual parallel sentences of unknown formality. We introduce a novel training scheme for multi-task models that automatically generates synthetic training triplets by inferring the missing element on the fly, thus enabling end-to-end training. Comprehensive automatic and human assessments show that our best model outperforms existing models by producing translations that better match desired formality levels while preserving the source meaning.1 Xing Niu 0001, Marine Carpuat |
AAAI | 1 |
| 2020 | Evaluating Robustness to Input Perturbations for Neural Machine TranslationabstractNeural Machine Translation (NMT) models are sensitive to small perturbations in the input.Robustness to such perturbations is typically measured using translation quality metrics such as BLEU on the noisy input.This paper proposes additional metrics which measure the relative degradation and changes in translation when small perturbations are added to the input.We focus on a class of models employing subword regularization to address robustness and perform extensive evaluations of these models using the robustness measures proposed.Results show that our proposed metrics reveal a clear trend of improved robustness to perturbations when subword regularization methods are used. Xing Niu 0001, Prashant Mathur, Georgiana Dinu, Yaser Al-Onaizan |
ACL | 1 |
| 2020 | Knowledge graph construction from multiple online encyclopedias
Tianxing Wu 0001, Haofen Wang, Guilin Qi, Xing Niu 0001, Meng Wang 0009, Chaomin Shi |
World Wide Web | 5 |
| 2018 | Multi-Task Neural Models for Translating Between Styles Within and Across LanguagesabstractGenerating natural language requires conveying content in an appropriate style. We explore two related tasks on generating text of varying formality: monolingual formality transfer and formality-sensitive machine translation. We propose to solve these tasks jointly using multi-task learning, and show that our models achieve state-of-the-art performance for formality transfer and are able to perform formality-sensitive translation without being explicitly trained on style-annotated translation examples. Xing Niu 0001, Sudha Rao, Marine Carpuat |
COLING | 1 |
| 2018 | Identifying Semantic Divergences in Parallel Text without AnnotationsabstractYogarshi Vyas, Xing Niu, Marine Carpuat. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Yogarshi Vyas, Xing Niu 0001, Marine Carpuat |
NAACL-HLT | 2 |
| 2017 | A Study of Style in Machine Translation: Controlling the Formality of Machine Translation OutputabstractStylistic variations of language, such as formality, carry speakers' intention beyond literal meaning and should be conveyed adequately in translation.We propose to use lexical formality models to control the formality level of machine translation output.We demonstrate the effectiveness of our approach in empirical evaluations, as measured by automatic metrics and human assessments. Xing Niu 0001, Marianna J. Martindale, Marine Carpuat |
EMNLP | 1 |
| 2016 | Compressing and Decoding Term Statistics Time Series
Jinfeng Rao, Xing Niu 0001, Jimmy Lin |
ECIR | 2 |
| 2012 | An effective rule miner for instance matching in a web of dataabstractPublishing structured data and linking them to Linking Open Data (LOD) is an ongoing effort to create a Web of data. Each newly involved data source may contain duplicated instances (entities) whose descriptions or schemata differ from those of the existing sources in LOD. To tackle this heterogeneity issue, several matching methods have been developed to link equivalent entities together. Many general-purpose matching methods which focus on similarity metrics suffer from very diverse matching results for different data source pairs. On the other hand, the dataset-specific ones leverage heuristic rules or even manual efforts to ensure the quality, which makes it impossible to apply them to other sources or domains. In this paper, we offer a third choice, a general method of automatically discovering dataset-specific matching rules. In particular, we propose a semi-supervised learning algorithm to iteratively refine matching rules and find new matches of high confidence based on these rules. This dramatically relieves the burden on users of defining rules but still gives high-quality matching results. We carry out experiments on real-world large scale data sources in LOD; the results show the effectiveness of our approach in terms of the precision of discovered matches and the number of missing matches found. Furthermore, we discuss several extensions (like similarity embedded rules, class restriction and SPARQL rewriting) to fit various applications with different requirements. Xing Niu 0001, Shu Rong, Haofen Wang, Yong Yu 0001 |
CIKM | 1 |
| 2012 | A Machine Learning Approach for Instance Matching Based on Similarity Metrics
Shu Rong, Xing Niu 0001, Evan Wei Xiang, Haofen Wang, Qiang Yang 0001, Yong Yu 0001 |
ISWC (1) | 2 |
| 2011 | Evaluating the Stability and Credibility of Ontology Matching Methods
Xing Niu 0001, Haofen Wang, Gang Wu 0007, Guilin Qi, Yong Yu 0001 |
ESWC (1) | 1 |
| 2011 | Zhishi.me - Weaving Chinese Linking Open Data
Xing Niu 0001, Xinruo Sun, Haofen Wang, Shu Rong, Guilin Qi, Yong Yu 0001 |
ISWC (2) | 1 |