VLDB 2026 Research / reviewers in the wild / expert
Ilia Kulikov
dblp:200/0187 · also Ilya Kulikov
· DBLP profile ↗
16ranked-venue papers
1as first author
10since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Language models and text generation · 48% Machine translation · 23% Efficient and distributed learning · 7% |
Topics — the 18 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
instruction following |
0.9 | 1 | 2025 | Following Length Constraints in Instructions · EMNLP 2025 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.9 | 1 | 2025 | NaturalReasoning: Reasoning in the Wild with 2.8M Challenging Questions · NeurIPS 2025 |
Natural language and speech › Language models and text generation
unlikelihood training |
0.9 | 2 | 2020 | Neural Text Generation With Unlikelihood Training · ICLR 2020 Don't Say That! Making Inconsistent Dialogue Unlikely with Unlikelihood Training · ACL 2020 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › masked modeling
masked prediction |
0.8 | 1 | 2024 | Multi-resolution HuBERT: Multi-resolution Speech Self-Supervised Learning with Masked Unit Prediction · ICLR 2024 |
Natural language and speech › Speech recognition and synthesis
speech representation learning |
0.8 | 1 | 2024 | Multi-resolution HuBERT: Multi-resolution Speech Self-Supervised Learning with Masked Unit Prediction · ICLR 2024 |
Natural language and speech › Machine translation › speech translation
direct speech translation |
0.7 | 1 | 2023 | UnitY: Two-pass Direct Speech-to-speech Translation with Discrete Units · ACL (1) 2023 |
Natural language and speech › Machine translation › speech translation
speech-to-speech translation |
0.7 | 1 | 2023 | UnitY: Two-pass Direct Speech-to-speech Translation with Discrete Units · ACL (1) 2023 |
Natural language and speech › Machine translation
speech translation |
0.7 | 1 | 2023 | Simple and Effective Unsupervised Speech Translation · ACL (1) 2023 |
Natural language and speech › Machine translation
unsupervised machine translation |
0.7 | 1 | 2023 | Simple and Effective Unsupervised Speech Translation · ACL (1) 2023 |
Natural language and speech › Language models and text generation › decoding › decoding strategy
beam search |
0.6 | 1 | 2022 | Uncertainty Determines the Adequacy of the Mode and the Tractability of Decoding in Sequence-to-Sequence Models · ACL (1) 2022 |
Natural language and speech › Language models and text generation › language modeling › language model architecture
sequence-to-sequence model |
0.6 | 1 | 2022 | Uncertainty Determines the Adequacy of the Mode and the Tractability of Decoding in Sequence-to-Sequence Models · ACL (1) 2022 |
Natural language and speech › Language models and text generation › decoding
decoding strategy |
0.4 | 1 | 2020 | Consistency of a Recurrent Language Model With Respect to Incomplete Decoding · EMNLP (1) 2020 |
Natural language and speech › Question answering and dialogue systems
dialogue generation |
0.4 | 1 | 2020 | Don't Say That! Making Inconsistent Dialogue Unlikely with Unlikelihood Training · ACL 2020 |
Natural language and speech › Language models and text generation › text generation
neural text generation |
0.4 | 1 | 2020 | Neural Text Generation With Unlikelihood Training · ICLR 2020 |
Natural language and speech › Language models and text generation › decoding
nucleus sampling |
0.4 | 1 | 2020 | Consistency of a Recurrent Language Model With Respect to Incomplete Decoding · EMNLP (1) 2020 |
Machine learning › Deep learning architectures and training
training objective |
0.4 | 1 | 2020 | Don't Say That! Making Inconsistent Dialogue Unlikely with Unlikelihood Training · ACL 2020 |
Machine learning › Transfer learning and domain adaptation › domain adaptation › unsupervised domain adaptation
self-training |
0.3 | 1 | 2025 | NaturalReasoning: Reasoning in the Wild with 2.8M Challenging Questions · NeurIPS 2025 |
Natural language and speech › Machine translation
neural machine translation |
0.2 | 1 | 2022 | Uncertainty Determines the Adequacy of the Mode and the Tractability of Decoding in Sequence-to-Sequence Models · ACL (1) 2022 |
Methods — techniques the papers use, named apart from their topics
self-rewarding · 0.9knowledge distillation · 0.9instruction tuning · 0.9masked unit prediction · 0.8hierarchical transformer · 0.8unsupervised learning · 0.7two-pass decoding · 0.7self-supervised learning · 0.7discrete units · 0.7exact n-best search · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Following Length Constraints in InstructionsabstractAligned instruction following models can better fulfill user requests than their unaligned counterparts.However, it has been shown that there is a length bias in evaluation of such models, and that training algorithms tend to exploit this bias by learning longer responses.In this work we show how to train models that can be controlled at inference time with instructions containing desired length constraints.Such models are superior in length instructed evaluations, outperforming standard instruction following models such as GPT4, Llama 3 and Mixtral. Weizhe Yuan, Ilia Kulikov, Kyunghyun Cho, Sainbayar Sukhbaatar, Jason Weston, Jing Xu 0014 |
EMNLP | 2 |
| 2025 | NaturalReasoning: Reasoning in the Wild with 2.8M Challenging QuestionsabstractScaling reasoning capabilities beyond traditional domains such as math and coding is hindered by the lack of diverse and high-quality questions. To overcome this limitation, we introduce a scalable approach for generating diverse and challenging reasoning questions, accompanied by reference answers. We present NaturalReasoning, a comprehensive dataset comprising 2.8 million questions that span multiple domains, including STEM fields (e.g., Physics, Computer Science), Economics, Social Sciences, and more. We demonstrate the utility of the questions in NaturalReasoning through knowledge distillation experiments which show that NaturalReasoning can effectively elicit and transfer reasoning capabilities from a strong teacher model. Furthermore, we demonstrate that NaturalReasoning is also effective for unsupervised self-training using external reward models or self-rewarding. Weizhe Yuan, Jane Dwivedi-Yu, Karthik Padthe, Ilia Kulikov, Kyunghyun Cho, Yuandong Tian, Jason Weston, Xian Li 0003 |
NeurIPS | 7 |
| 2024 | Multi-resolution HuBERT: Multi-resolution Speech Self-Supervised Learning with Masked Unit PredictionabstractExisting Self-Supervised Learning (SSL) models for speech typically process speech signals at a fixed resolution of 20 milliseconds. This approach overlooks the varying informational content present at different resolutions in speech signals. In contrast, this paper aims to incorporate multi-resolution information into speech self-supervised representation learning. We introduce an SSL model that leverages a hierarchical Transformer architecture, complemented by HuBERT-style masked prediction objectives, to process speech at multiple resolutions. Experimental results indicate that the proposed model not only achieves more efficient inference but also exhibits superior or comparable performance to the original HuBERT model over various tasks. Specifically, significant performance improvements over the original HuBERT have been observed in fine-tuning experiments on the LibriSpeech speech recognition benchmark as well as in evaluations using the Speech Universal PERformance Benchmark (SUPERB) and Multilingual SUPERB (ML-SUPERB). Jiatong Shi, Hirofumi Inaguma, Xutai Ma, Ilia Kulikov, Anna Y. Sun |
ICLR | 4 |
| 2024 | Investigating Decoder-only Large Language Models for Speech-to-text Translation
Chao-Wei Huang, Hongyu Gong, Hirofumi Inaguma, Ilia Kulikov, Ruslan Mavlyutov, Sravya Popuri |
INTERSPEECH | 5 |
| 2024 | Massively Multilingual Forced Aligner Leveraging Self-Supervised Discrete UnitsabstractWe propose a massively multilingual speech-to-text neural forced aligner that supports 98 languages with a single architecture. The aligner takes self-supervised discrete acoustic units and unnormalized characters including punctuation marks as inputs. We train the aligner as a part of a non-autoregressive text-to-unit (T2U) model without any external aligner. The T2U model is trained on speech-text paired data in various domains and recording conditions. Experimental evaluation demonstrates that the proposed T2U aligner achieves competitive quality to existing monolingual aligners while supporting much more languages. We also showcase a zero-shot forced alignment capability on unseen languages. Hirofumi Inaguma, Ilia Kulikov, Zhaoheng Ni, Sravya Popuri, Paden Tomasello |
SLT | 2 |
| 2023 | UnitY: Two-pass Direct Speech-to-speech Translation with Discrete UnitsabstractHirofumi Inaguma, Sravya Popuri, Ilia Kulikov, Peng-Jen Chen, Changhan Wang, Yu-An Chung, Yun Tang, Ann Lee, Shinji Watanabe, Juan Pino. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Hirofumi Inaguma, Sravya Popuri, Ilia Kulikov, Peng-Jen Chen, Changhan Wang, Yu-An Chung, Yun Tang 0002, Ann Lee 0001, Shinji Watanabe 0001, Juan Pino 0001 |
ACL (1) | 3 |
| 2023 | Simple and Effective Unsupervised Speech TranslationabstractChanghan Wang, Hirofumi Inaguma, Peng-Jen Chen, Ilia Kulikov, Yun Tang, Wei-Ning Hsu, Michael Auli, Juan Pino. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Changhan Wang, Hirofumi Inaguma, Peng-Jen Chen, Ilia Kulikov, Yun Tang 0002, Wei-Ning Hsu, Michael Auli, Juan Pino 0001 |
ACL (1) | 4 |
| 2023 | Named Entity Detection and Injection for Direct Speech TranslationabstractIn a sentence, certain words are critical for its semantic. Among them, named entities (NEs) are notoriously challenging for neural models. Despite their importance, their accurate handling has been neglected in speech-to-text (S2T) translation research, and recent work has shown that S2T models perform poorly for locations and notably person names, whose spelling is challenging unless known in advance. In this work, we explore how to leverage dictionaries of NEs known to likely appear in a given context to improve S2T model outputs. Our experiments show that we can reliably detect NEs likely present in an utterance starting from S2T encoder outputs. Indeed, we demonstrate that the current detection quality is sufficient to improve NE accuracy in the translation with a 31% reduction in person name errors. Marco Gaido, Yun Tang 0002, Ilia Kulikov, Rongqing Huang, Hongyu Gong, Hirofumi Inaguma |
ICASSP | 3 |
| 2023 | Improving Speech-to-Speech Translation Through Unlabeled TextabstractDirect speech-to-speech translation (S2ST) is among the most challenging problems in the translation paradigm due to the significant scarcity of S2ST data. While effort has been made to increase the data size from unlabeled speech by cascading pretrained speech recognition (ASR), machine translation (MT) and text-to-speech (TTS) models; unlabeled text has remained relatively under-utilized to improve S2ST. We propose an effective way to utilize the massive existing unlabeled text from different languages to create a large amount of S2ST data to improve S2ST performance by applying various acoustic effects to the generated synthetic data. Empirically our method outperforms the state of the art in Spanish-English translation by up to 2 BLEU. Significant gains by the proposed method are demonstrated in extremely low-resource settings for both Spanish-English and Russian-English translations. Xuan-Phi Nguyen, Sravya Popuri, Changhan Wang, Yun Tang 0002, Ilia Kulikov, Hongyu Gong |
ICASSP | 5 |
| 2022 | Uncertainty Determines the Adequacy of the Mode and the Tractability of Decoding in Sequence-to-Sequence ModelsabstractIn many natural language processing (NLP) tasks the same input (e.g.source sentence) can have multiple possible outputs (e.g.translations).To analyze how this ambiguity (also known as intrinsic uncertainty) shapes the distribution learned by neural sequence models we measure sentence-level uncertainty by computing the degree of overlap between references in multi-reference test sets from two different NLP tasks: machine translation (MT) and grammatical error correction (GEC).At both the sentence-and the task-level, intrinsic uncertainty has major implications for various aspects of search such as the inductive biases in beam search and the complexity of exact search.In particular, we show that well-known pathologies such as a high number of beam search errors, the inadequacy of the mode, and the drop in system performance with large beam sizes apply to tasks with high level of ambiguity such as MT but not to less uncertain tasks such as GEC.Furthermore, we propose a novel exact n-best search algorithm for neural sequence models, and show that intrinsic uncertainty affects model uncertainty as the model tends to overly spread out the probability mass for uncertain tasks and sentences. Felix Stahlberg, Ilia Kulikov, Shankar Kumar |
ACL (1) | 2 |
| 2020 | Don't Say That! Making Inconsistent Dialogue Unlikely with Unlikelihood TrainingabstractGenerative dialogue models currently suffer from a number of problems which standard maximum likelihood training does not address.They tend to produce generations that (i) rely too much on copying from the context, (ii) contain repetitions within utterances, (iii) overuse frequent words, and (iv) at a deeper level, contain logical flaws.In this work we show how all of these problems can be addressed by extending the recently introduced unlikelihood loss (Welleck et al., 2019a) to these cases.We show that appropriate loss functions which regularize generated outputs to match human distributions are effective for the first three issues.For the last important general issue, we show applying unlikelihood to collected data of what a model should not do is effective for improving logical consistency, potentially paving the way to generative models with greater reasoning ability.We demonstrate the efficacy of our approach across several dialogue tasks. Margaret Li, Stephen Roller, Ilia Kulikov, Sean Welleck, Y-Lan Boureau, Kyunghyun Cho, Jason Weston |
ACL | 3 |
| 2020 | Consistency of a Recurrent Language Model With Respect to Incomplete DecodingabstractDespite strong performance on a variety of tasks, neural sequence models trained with maximum likelihood have been shown to exhibit issues such as length bias and degenerate repetition.We study the related issue of receiving infinite-length sequences from a recurrent language model when using common decoding algorithms.To analyze this issue, we first define inconsistency of a decoding algorithm, meaning that the algorithm can yield an infinite-length sequence that has zero probability under the model.We prove that commonly used incomplete decoding algorithms -greedy search, beam search, top-k sampling, and nucleus sampling -are inconsistent, despite the fact that recurrent language models are trained to produce sequences of finite length.Based on these insights, we propose two remedies which address inconsistency: consistent variants of top-k and nucleus sampling, and a selfterminating recurrent language model.Empirical results show that inconsistency occurs in practice, and that the proposed methods prevent inconsistency. Sean Welleck, Ilia Kulikov, Jaedeok Kim, Richard Yuanzhe Pang, Kyunghyun Cho |
EMNLP (1) | 2 |
| 2020 | Neural Text Generation With Unlikelihood Training
Sean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan, Kyunghyun Cho, Jason Weston |
ICLR | 2 |
| 2019 | Importance of Search and Evaluation Strategies in Neural Dialogue ModelingabstractWe investigate the impact of search strategies in neural dialogue modeling.We first compare two standard search algorithms, greedy and beam search, as well as our newly proposed iterative beam search which produces a more diverse set of candidate responses.We evaluate these strategies in realistic full conversations with humans and propose a modelbased Bayesian calibration to address annotator bias.These conversations are analyzed using two automatic metrics: log-probabilities assigned by the model and utterance diversity.Our experiments reveal that better search algorithms lead to higher rated conversations.However, finding the optimal selection mechanism to choose from a more diverse set of candidates is still an open question. Ilia Kulikov, Alexander H. Miller, Kyunghyun Cho, Jason Weston |
INLG | 1 |
| 2017 | Returnn: The RWTH extensible training framework for universal recurrent neural networksabstractIn this work we release our extensible and easily configurable neural network training software. It provides a rich set of functional layers with a particular focus on efficient training of recurrent neural network topologies on multiple GPUs. The source of the software package is public and freely available for academic research purposes and can be used as a framework or as a standalone tool which supports a flexible configuration. The software allows to train state-of-the-art deep bidirectional long short-term memory (LSTM) models on both one dimensional data like speech or two dimensional data like handwritten text and was used to develop successful submission systems in several evaluation campaigns. Patrick Doetsch, Albert Zeyer, Paul Voigtlaender, Ilia Kulikov, Ralf Schlüter, Hermann Ney |
ICASSP | 4 |
| 2017 | Faster sequence trainingabstractIt has been shown that sequence-discriminative training can improve the performance for large vocabulary continuous speech recognition. Our main contribution is a novel method for reducing the computation time of any sort of sequence training while only slightly decreasing the overall performance. The method allows to parallelize the forward propagation through the network, the loss and loss gradient calculation which will provide a frame-wise error signal, and an independent forward and back propagation using that error signal. That last step can be calculated in a frame-wise manner and thus allows to use frame chunking to further improve the runtime. The loss calculation can itself be parallelized over many sequences. In addition to several experiments which outline the runtime gains, we also provide a convergence proof sketch. We extend on the research of sequence training of bidirectional long-short term memory ((B)LSTM) networks and provide an overview and comparison over different criteria. We have published all the code as part of our RETURNN and RASR framework including our training setup configurations. Albert Zeyer, Ilia Kulikov, Ralf Schlüter, Hermann Ney |
ICASSP | 2 |