Ilia Kulikov

dblp:200/0187 · also Ilya Kulikov · DBLP profile ↗
← Back
16ranked-venue papers
1as first author
10since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Language models and text generation · 48% Machine translation · 23% Efficient and distributed learning · 7%

Topics — the 18 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
instruction following
0.912025
Following Length Constraints in Instructions · EMNLP 2025
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.912025
NaturalReasoning: Reasoning in the Wild with 2.8M Challenging Questions · NeurIPS 2025
Natural language and speech › Language models and text generation
unlikelihood training
0.922020
Neural Text Generation With Unlikelihood Training · ICLR 2020
Don't Say That! Making Inconsistent Dialogue Unlikely with Unlikelihood Training · ACL 2020
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › masked modeling
masked prediction
0.812024
Multi-resolution HuBERT: Multi-resolution Speech Self-Supervised Learning with Masked Unit Prediction · ICLR 2024
Natural language and speech › Speech recognition and synthesis
speech representation learning
0.812024
Multi-resolution HuBERT: Multi-resolution Speech Self-Supervised Learning with Masked Unit Prediction · ICLR 2024
Natural language and speech › Machine translation › speech translation
direct speech translation
0.712023
UnitY: Two-pass Direct Speech-to-speech Translation with Discrete Units · ACL (1) 2023
Natural language and speech › Machine translation › speech translation
speech-to-speech translation
0.712023
UnitY: Two-pass Direct Speech-to-speech Translation with Discrete Units · ACL (1) 2023
Natural language and speech › Machine translation
speech translation
0.712023
Simple and Effective Unsupervised Speech Translation · ACL (1) 2023
Natural language and speech › Machine translation
unsupervised machine translation
0.712023
Simple and Effective Unsupervised Speech Translation · ACL (1) 2023
Natural language and speech › Language models and text generation › decoding › decoding strategy
beam search
0.612022
Uncertainty Determines the Adequacy of the Mode and the Tractability of Decoding in Sequence-to-Sequence Models · ACL (1) 2022
Natural language and speech › Language models and text generation › language modeling › language model architecture
sequence-to-sequence model
0.612022
Uncertainty Determines the Adequacy of the Mode and the Tractability of Decoding in Sequence-to-Sequence Models · ACL (1) 2022
Natural language and speech › Language models and text generation › decoding
decoding strategy
0.412020
Consistency of a Recurrent Language Model With Respect to Incomplete Decoding · EMNLP (1) 2020
Natural language and speech › Question answering and dialogue systems
dialogue generation
0.412020
Don't Say That! Making Inconsistent Dialogue Unlikely with Unlikelihood Training · ACL 2020
Natural language and speech › Language models and text generation › text generation
neural text generation
0.412020
Neural Text Generation With Unlikelihood Training · ICLR 2020
Natural language and speech › Language models and text generation › decoding
nucleus sampling
0.412020
Consistency of a Recurrent Language Model With Respect to Incomplete Decoding · EMNLP (1) 2020
Machine learning › Deep learning architectures and training
training objective
0.412020
Don't Say That! Making Inconsistent Dialogue Unlikely with Unlikelihood Training · ACL 2020
Machine learning › Transfer learning and domain adaptation › domain adaptation › unsupervised domain adaptation
self-training
0.312025
NaturalReasoning: Reasoning in the Wild with 2.8M Challenging Questions · NeurIPS 2025
Natural language and speech › Machine translation
neural machine translation
0.212022
Uncertainty Determines the Adequacy of the Mode and the Tractability of Decoding in Sequence-to-Sequence Models · ACL (1) 2022

Methods — techniques the papers use, named apart from their topics

self-rewarding · 0.9knowledge distillation · 0.9instruction tuning · 0.9masked unit prediction · 0.8hierarchical transformer · 0.8unsupervised learning · 0.7two-pass decoding · 0.7self-supervised learning · 0.7discrete units · 0.7exact n-best search · 0.6
YearPublicationVenuePosition
2025 Following Length Constraints in Instructions
abstract
Aligned instruction following models can better fulfill user requests than their unaligned counterparts.However, it has been shown that there is a length bias in evaluation of such models, and that training algorithms tend to exploit this bias by learning longer responses.In this work we show how to train models that can be controlled at inference time with instructions containing desired length constraints.Such models are superior in length instructed evaluations, outperforming standard instruction following models such as GPT4, Llama 3 and Mixtral.
Weizhe Yuan, Ilia Kulikov, Kyunghyun Cho, Sainbayar Sukhbaatar, Jason Weston, Jing Xu 0014
EMNLP2
2025 NaturalReasoning: Reasoning in the Wild with 2.8M Challenging Questions
abstract
Scaling reasoning capabilities beyond traditional domains such as math and coding is hindered by the lack of diverse and high-quality questions. To overcome this limitation, we introduce a scalable approach for generating diverse and challenging reasoning questions, accompanied by reference answers. We present NaturalReasoning, a comprehensive dataset comprising 2.8 million questions that span multiple domains, including STEM fields (e.g., Physics, Computer Science), Economics, Social Sciences, and more. We demonstrate the utility of the questions in NaturalReasoning through knowledge distillation experiments which show that NaturalReasoning can effectively elicit and transfer reasoning capabilities from a strong teacher model. Furthermore, we demonstrate that NaturalReasoning is also effective for unsupervised self-training using external reward models or self-rewarding.
Weizhe Yuan, Jane Dwivedi-Yu, Karthik Padthe, Ilia Kulikov, Kyunghyun Cho, Yuandong Tian, Jason Weston, Xian Li 0003
NeurIPS7
2024 Multi-resolution HuBERT: Multi-resolution Speech Self-Supervised Learning with Masked Unit Prediction
abstract
Existing Self-Supervised Learning (SSL) models for speech typically process speech signals at a fixed resolution of 20 milliseconds. This approach overlooks the varying informational content present at different resolutions in speech signals. In contrast, this paper aims to incorporate multi-resolution information into speech self-supervised representation learning. We introduce an SSL model that leverages a hierarchical Transformer architecture, complemented by HuBERT-style masked prediction objectives, to process speech at multiple resolutions. Experimental results indicate that the proposed model not only achieves more efficient inference but also exhibits superior or comparable performance to the original HuBERT model over various tasks. Specifically, significant performance improvements over the original HuBERT have been observed in fine-tuning experiments on the LibriSpeech speech recognition benchmark as well as in evaluations using the Speech Universal PERformance Benchmark (SUPERB) and Multilingual SUPERB (ML-SUPERB).
Jiatong Shi, Hirofumi Inaguma, Xutai Ma, Ilia Kulikov, Anna Y. Sun
ICLR4
2024 Investigating Decoder-only Large Language Models for Speech-to-text Translation
Chao-Wei Huang, Hongyu Gong, Hirofumi Inaguma, Ilia Kulikov, Ruslan Mavlyutov, Sravya Popuri
INTERSPEECH5
2024 Massively Multilingual Forced Aligner Leveraging Self-Supervised Discrete Units
abstract
We propose a massively multilingual speech-to-text neural forced aligner that supports 98 languages with a single architecture. The aligner takes self-supervised discrete acoustic units and unnormalized characters including punctuation marks as inputs. We train the aligner as a part of a non-autoregressive text-to-unit (T2U) model without any external aligner. The T2U model is trained on speech-text paired data in various domains and recording conditions. Experimental evaluation demonstrates that the proposed T2U aligner achieves competitive quality to existing monolingual aligners while supporting much more languages. We also showcase a zero-shot forced alignment capability on unseen languages.
Hirofumi Inaguma, Ilia Kulikov, Zhaoheng Ni, Sravya Popuri, Paden Tomasello
SLT2
2023 UnitY: Two-pass Direct Speech-to-speech Translation with Discrete Units
abstract
Hirofumi Inaguma, Sravya Popuri, Ilia Kulikov, Peng-Jen Chen, Changhan Wang, Yu-An Chung, Yun Tang, Ann Lee, Shinji Watanabe, Juan Pino. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Hirofumi Inaguma, Sravya Popuri, Ilia Kulikov, Peng-Jen Chen, Changhan Wang, Yu-An Chung, Yun Tang 0002, Ann Lee 0001, Shinji Watanabe 0001, Juan Pino 0001
ACL (1)3
2023 Simple and Effective Unsupervised Speech Translation
abstract
Changhan Wang, Hirofumi Inaguma, Peng-Jen Chen, Ilia Kulikov, Yun Tang, Wei-Ning Hsu, Michael Auli, Juan Pino. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Changhan Wang, Hirofumi Inaguma, Peng-Jen Chen, Ilia Kulikov, Yun Tang 0002, Wei-Ning Hsu, Michael Auli, Juan Pino 0001
ACL (1)4
2023 Named Entity Detection and Injection for Direct Speech Translation
abstract
In a sentence, certain words are critical for its semantic. Among them, named entities (NEs) are notoriously challenging for neural models. Despite their importance, their accurate handling has been neglected in speech-to-text (S2T) translation research, and recent work has shown that S2T models perform poorly for locations and notably person names, whose spelling is challenging unless known in advance. In this work, we explore how to leverage dictionaries of NEs known to likely appear in a given context to improve S2T model outputs. Our experiments show that we can reliably detect NEs likely present in an utterance starting from S2T encoder outputs. Indeed, we demonstrate that the current detection quality is sufficient to improve NE accuracy in the translation with a 31% reduction in person name errors.
Marco Gaido, Yun Tang 0002, Ilia Kulikov, Rongqing Huang, Hongyu Gong, Hirofumi Inaguma
ICASSP3
2023 Improving Speech-to-Speech Translation Through Unlabeled Text
abstract
Direct speech-to-speech translation (S2ST) is among the most challenging problems in the translation paradigm due to the significant scarcity of S2ST data. While effort has been made to increase the data size from unlabeled speech by cascading pretrained speech recognition (ASR), machine translation (MT) and text-to-speech (TTS) models; unlabeled text has remained relatively under-utilized to improve S2ST. We propose an effective way to utilize the massive existing unlabeled text from different languages to create a large amount of S2ST data to improve S2ST performance by applying various acoustic effects to the generated synthetic data. Empirically our method outperforms the state of the art in Spanish-English translation by up to 2 BLEU. Significant gains by the proposed method are demonstrated in extremely low-resource settings for both Spanish-English and Russian-English translations.
Xuan-Phi Nguyen, Sravya Popuri, Changhan Wang, Yun Tang 0002, Ilia Kulikov, Hongyu Gong
ICASSP5
2022 Uncertainty Determines the Adequacy of the Mode and the Tractability of Decoding in Sequence-to-Sequence Models
abstract
In many natural language processing (NLP) tasks the same input (e.g.source sentence) can have multiple possible outputs (e.g.translations).To analyze how this ambiguity (also known as intrinsic uncertainty) shapes the distribution learned by neural sequence models we measure sentence-level uncertainty by computing the degree of overlap between references in multi-reference test sets from two different NLP tasks: machine translation (MT) and grammatical error correction (GEC).At both the sentence-and the task-level, intrinsic uncertainty has major implications for various aspects of search such as the inductive biases in beam search and the complexity of exact search.In particular, we show that well-known pathologies such as a high number of beam search errors, the inadequacy of the mode, and the drop in system performance with large beam sizes apply to tasks with high level of ambiguity such as MT but not to less uncertain tasks such as GEC.Furthermore, we propose a novel exact n-best search algorithm for neural sequence models, and show that intrinsic uncertainty affects model uncertainty as the model tends to overly spread out the probability mass for uncertain tasks and sentences.
Felix Stahlberg, Ilia Kulikov, Shankar Kumar
ACL (1)2
2020 Don't Say That! Making Inconsistent Dialogue Unlikely with Unlikelihood Training
abstract
Generative dialogue models currently suffer from a number of problems which standard maximum likelihood training does not address.They tend to produce generations that (i) rely too much on copying from the context, (ii) contain repetitions within utterances, (iii) overuse frequent words, and (iv) at a deeper level, contain logical flaws.In this work we show how all of these problems can be addressed by extending the recently introduced unlikelihood loss (Welleck et al., 2019a) to these cases.We show that appropriate loss functions which regularize generated outputs to match human distributions are effective for the first three issues.For the last important general issue, we show applying unlikelihood to collected data of what a model should not do is effective for improving logical consistency, potentially paving the way to generative models with greater reasoning ability.We demonstrate the efficacy of our approach across several dialogue tasks.
Margaret Li, Stephen Roller, Ilia Kulikov, Sean Welleck, Y-Lan Boureau, Kyunghyun Cho, Jason Weston
ACL3
2020 Consistency of a Recurrent Language Model With Respect to Incomplete Decoding
abstract
Despite strong performance on a variety of tasks, neural sequence models trained with maximum likelihood have been shown to exhibit issues such as length bias and degenerate repetition.We study the related issue of receiving infinite-length sequences from a recurrent language model when using common decoding algorithms.To analyze this issue, we first define inconsistency of a decoding algorithm, meaning that the algorithm can yield an infinite-length sequence that has zero probability under the model.We prove that commonly used incomplete decoding algorithms -greedy search, beam search, top-k sampling, and nucleus sampling -are inconsistent, despite the fact that recurrent language models are trained to produce sequences of finite length.Based on these insights, we propose two remedies which address inconsistency: consistent variants of top-k and nucleus sampling, and a selfterminating recurrent language model.Empirical results show that inconsistency occurs in practice, and that the proposed methods prevent inconsistency.
Sean Welleck, Ilia Kulikov, Jaedeok Kim, Richard Yuanzhe Pang, Kyunghyun Cho
EMNLP (1)2
2020 Neural Text Generation With Unlikelihood Training
Sean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan, Kyunghyun Cho, Jason Weston
ICLR2
2019 Importance of Search and Evaluation Strategies in Neural Dialogue Modeling
abstract
We investigate the impact of search strategies in neural dialogue modeling.We first compare two standard search algorithms, greedy and beam search, as well as our newly proposed iterative beam search which produces a more diverse set of candidate responses.We evaluate these strategies in realistic full conversations with humans and propose a modelbased Bayesian calibration to address annotator bias.These conversations are analyzed using two automatic metrics: log-probabilities assigned by the model and utterance diversity.Our experiments reveal that better search algorithms lead to higher rated conversations.However, finding the optimal selection mechanism to choose from a more diverse set of candidates is still an open question.
Ilia Kulikov, Alexander H. Miller, Kyunghyun Cho, Jason Weston
INLG1
2017 Returnn: The RWTH extensible training framework for universal recurrent neural networks
abstract
In this work we release our extensible and easily configurable neural network training software. It provides a rich set of functional layers with a particular focus on efficient training of recurrent neural network topologies on multiple GPUs. The source of the software package is public and freely available for academic research purposes and can be used as a framework or as a standalone tool which supports a flexible configuration. The software allows to train state-of-the-art deep bidirectional long short-term memory (LSTM) models on both one dimensional data like speech or two dimensional data like handwritten text and was used to develop successful submission systems in several evaluation campaigns.
Patrick Doetsch, Albert Zeyer, Paul Voigtlaender, Ilia Kulikov, Ralf Schlüter, Hermann Ney
ICASSP4
2017 Faster sequence training
abstract
It has been shown that sequence-discriminative training can improve the performance for large vocabulary continuous speech recognition. Our main contribution is a novel method for reducing the computation time of any sort of sequence training while only slightly decreasing the overall performance. The method allows to parallelize the forward propagation through the network, the loss and loss gradient calculation which will provide a frame-wise error signal, and an independent forward and back propagation using that error signal. That last step can be calculated in a frame-wise manner and thus allows to use frame chunking to further improve the runtime. The loss calculation can itself be parallelized over many sequences. In addition to several experiments which outline the runtime gains, we also provide a convergence proof sketch. We extend on the research of sequence training of bidirectional long-short term memory ((B)LSTM) networks and provide an overview and comparison over different criteria. We have published all the code as part of our RETURNN and RASR framework including our training setup configurations.
Albert Zeyer, Ilia Kulikov, Ralf Schlüter, Hermann Ney
ICASSP2